Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Oct 199319922001200920182026
48 results for practical implications

Study shows better dropout prediction from clickstream data in MOOCs.

problem Evaluating predictive models of student success in MOOCs.
method Statistical testing of hypotheses about model performance, focusing on algorithms and feature extraction methods.
result Clickstream-based feature extraction outperforms forum- and assignment-based methods in predicting student dropout.

Study predicts P2P lending platform failures using machine learning.

problem Predicting failures of P2P lending platforms in China.
method Used machine learning models with filter and wrapper methods, forward selection, and backward elimination.
result Identified robust variables for predicting platform failures with high AUC and F1 scores.

In manifold learning, algorithms based on graph Laplacians constructed from data have received considerable attention both in practical applications and theoretical analysis. In particular, the convergence of graph Laplacians obtained from sampled data to certain continuous operators has become an active research topic…

2011-05-19abs ↗pdf ↗

This paper closely examines theoretical and practical aspects of the widely used discounted cash flows (DCF) valuation method. It assesses its potentials as well as several weaknesses. A special emphasize is being put on the valuation of companies using the DCF method. The paper finds that the discounted cash flow meth…

2010-03-25abs ↗pdf ↗

Heterophily affects GNN robustness; separating ego- and neighbor-embeddings improves defense.

problem The robustness of GNNs to adversarial attacks.
method Formalized relation between heterophily and GNN robustness; empirical analysis; design principles for improved robustness.
result Separating ego- and neighbor-embeddings increases GNN robustness.

Most real life systems have a random component: the multitude of endogenous and exogenous factors influencing them result in stochastic fluctuations of the parameters determining their dynamics. These empirical systems are in many cases subject to noise of multiplicative nature. The special properties of multiplicative…

2008-07-11abs ↗pdf ↗

New method reconstructs significant parts of training data from neural networks.

problem Understanding and reconstructing training data from neural networks.
method Proposes a novel reconstruction scheme based on recent theoretical results about neural network training.
result Shows that a significant fraction of training data can be reconstructed from neural network parameters.

Paper investigates differentiable fuzzy implications and their suitability for learning.

problem Analyzing differentiable fuzzy implications and their suitability for learning.
method Investigates the properties of fuzzy implications in a differentiable setting and introduces a new family of fuzzy implications.
result Various fuzzy implications are unsuitable for differentiable learning, and a new family of fuzzy implications is introduced.

The paper develops no arbitrage results for trajectory based models by imposing general constraints on the trading portfolios. The main condition imposed, in order to avoid arbitrage opportunities, is a local continuity requirement on the final portfolio value considered as a functional on the trajectory space. The pap…

2014-03-22abs ↗pdf ↗

Study provides guarantees for kernel clustering under non-parametric mixtures.

problem Statistical guarantees for kernel-based clustering without strong assumptions.
method Non-parametric mixture models, kernel-based clustering, consistency guarantees.
result Necessary and sufficient separability conditions for consistent clustering recovery.

We attempt to explain stock market dynamics in terms of the interaction among three variables: market price, investor opinion and information flow. We propose a framework for such interaction and apply it to build a model of stock market dynamics which we study both empirically and theoretically. We demonstrate that th…

2014-04-29abs ↗pdf ↗

Pruning neural networks adds differential privacy noise, preserving data utility.

problem Achieving differential privacy in neural networks without sacrificing data utility.
method Proving equivalence between pruning and adding differential privacy noise to hidden-layer activations.
result Pruning can be a more effective alternative to adding differential privacy noise for neural networks.

TransINT embeds KGs by preserving implication rules, outperforming existing methods.

problem Embedding KGs while preserving relation implications for better access and analysis.
method Isomorphic intersections of linear subspaces with shared parameters for missing facts.
result Significant performance improvement in link prediction and triple classification.

Researchers prove Transformers are Turing-complete and analyze their components.

problem Understanding the computational power and limitations of Transformers.
method Analyzed Turing-completeness of vanilla and modified Transformers, and necessity of components.
result Transformers with positional masking and positional encodings are Turing-complete.

Paper improves sample efficiency of transfer learning in diffusion models.

problem Diffusion models need too much data to train from scratch.
method Assumes shared low-dimensional representation across tasks for improved sample efficiency.
result Sample complexity of target tasks can be reduced with a well-learned representation.

ProbE model improves relational implication detection to 0.8143.

problem Improving inference of relational data to extract more useful information.
method Formal probabilistic model of relational implication using estimators based on empirical distribution.
result ProbE model outperforms existing approaches, achieving 0.8143 AUC on evaluation dataset.

Normal surface theory is a central tool in algorithmic three-dimensional topology, and the enumeration of vertex normal surfaces is the computational bottleneck in many important algorithms. However, it is not well understood how the number of such surfaces grows in relation to the size of the underlying triangulation.…

2009-11-30abs ↗pdf ↗

The benefits of portfolio diversification is a central tenet implicit to modern financial theory and practice. Linked to diversification is the notion of breadth. Breadth is correctly thought of as the number of in- dependent bets available to an investor. Conventionally applications us- ing breadth frequently assume o…

2006-01-23abs ↗pdf ↗

Paper summarizes unsupervised learning challenges for disentangled representations.

problem Unsupervised learning of disentangled representations without inductive biases.
method Theoretical and practical analysis of existing approaches.
result Unsupervised disentanglement is fundamentally impossible without inductive biases.

In recent years, the spectral analysis of appropriately defined kernel matrices has emerged as a principled way to extract the low-dimensional structure often prevalent in high-dimensional data. Here we provide an introduction to spectral methods for linear and nonlinear dimension reduction, emphasizing ways to overcom…

2009-06-24abs ↗pdf ↗

Aggregation defenses improve deep learning models' robustness against data poisoning attacks.

problem Data poisoning attacks manipulate deep learning models with malicious training samples.
method Deep Partition Aggregation, efficiency improvements, data-to-complexity ratio, poisoning overfitting phenomenon.
result Aggregation defenses boost poisoning robustness through the poisoning overfitting phenomenon.

This paper uses neural networks to accurately model competing risks in survival analysis.

problem Ignoring competing risks leads to biased survival estimation in machine learning models.
method The paper introduces constrained monotonic neural networks to model each competing survival distribution.
result The method ensures exact likelihood maximization with reduced computational cost.

We describe many vantage points on the Baire metric and its use in clustering data, or its use in preprocessing and structuring data in order to support search and retrieval operations. In some cases, we proceed directly to clusters and do not directly determine the distances. We show how a hierarchical clustering can …

2011-11-27abs ↗pdf ↗

The aim of this paper is to introduce a method for computing the allocated Solvency II Capital Requirement (SCR) of each Risk which the company is exposed to, taking in account for the diversification effect among different risks. The method suggested is based on the Euler principle. We show that it has very suitable p…

2015-11-09abs ↗pdf ↗

Bayesian models overestimate clusters, but practical summaries can correct this.

problem Bayesian mixture models overestimate the number of clusters.
method Simulations and gene expression data analysis using MCMC summarisation.
result Overestimation is limited in finite samples and can be corrected, but misspecification leads to significant overestimation.

VolTS uses stats & ML to forecast stock market trends based on volatility.

problem Capturing profitable trading opportunities from market dynamics.
method Combines statistical analysis with machine learning; k-means++ clustering, Granger causality test.
result Effective at identifying profitable trading opportunities through volatility clusters and Granger causality.

Extends neural network approximation to probability measures and tree-structured data.

problem Universal approximation of functions on probability measures and tree-structured domains.
method Proof of neural network density in probability measure spaces and Cartesian products.
result Universal approximation theorem for tree-structured domains, including JSON.

Modern machine learning practices contradict traditional bias-variance theory.

problem Modern machine learning models often fit data perfectly, yet perform well.
method Introducing a 'double descent' curve to reconcile classical and modern practices.
result Increasing model capacity beyond interpolation improves performance.

The paper explores how to extrapolate from limited data points using causal mechanisms.

problem Handling distribution shifts with limited target samples.
method Formulates the extrapolation problem with a latent-variable model embodying the minimal change principle in causal mechanisms, and identifies conditions for identification.
result Theoretical understanding and practical methods for extrapolation without requiring an on-support target distribution.

The cycling operation endows the super summit set SxS_x of any element xx of a Garside group GG with the structure of a directed graph ΓxΓ_x. We establish that the subset UxU_x of SxS_x consisting of the circuits of ΓxΓ_x can be used instead of SxS_x for deciding conjugacy to xx in GG, yielding a faster and more pr…

2003-06-12abs ↗pdf ↗

Study derives new equation for reserves in non-monotone information scenarios.

problem Modeling reserves in situations where information is not always increasing.
method Infinitesimal approach to derive generalized stochastic Thiele equation.
result New equation allows for information discarding and solves open problems.