A form of generalisation error known as Off Training Set (OTS) error was recently introduced in [Wolpert, 1996b], along with a theorem showing that small training set error does not guarantee small OTS error, unless assumptions are made about the target function. Here it is shown that the applicability of this theorem …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The study analyzes and mitigates errors in PC-based causal discovery methods.
This work finds a point with small test error in polynomial time for mildly overparameterized neural nets.
Dynamic hedging of an European option under a general local volatility model with small linear transaction costs is studied. A continuous control version of Leland's strategy that asymptotically replicates the payoff is constructed. An associated central limit theorem of hedging error is proved. The asymptotic error va…
Langevin dynamics fails to produce accurate samples even with small score function errors.
Novel approach for SEM in small samples with .
Paper introduces a new method for error estimation in classification tasks with limited data.
New learner achieves optimal agnostic error in small error regime.
The pricing of options in exponential Levy models amounts to the computation of expectations of functionals of Levy processes. In many situations, Monte-Carlo methods are used. However, the simulation of a Levy process with infinite Levy measure generally requires either to truncate small jumps or to replace them by a …
Paper develops error rates for physics-informed learning, comparing it to data-driven methods.
The skip-connections used in residual networks have become a standard architecture choice in deep learning due to the increased training stability and generalization performance with this architecture, although there has been limited theoretical understanding for this improvement. In this work, we analyze overparameter…
New findings show score matching's accuracy doesn't ensure numerical stability in diffusion sampling.
Confidence measures for the generalization error are crucial when small training samples are used to construct classifiers. A common approach is to estimate the generalization error by resampling and then assume the resampled estimator follows a known distribution to form a confidence set [Kohavi 1995, Martin 1996,Yang…
ASGD outperforms SGD in overparameterized linear regression, especially in subspaces of small eigenvalues.
The paper gives a constructive method, based on greedy algorithms, that provides for the classes of functions with small mixed smoothness the best possible in the sense of order approximation error for the -term approximation with respect to the trigonometric system.
In this paper we define a small variation of the Taylor method and a formula for the global error of this new numerical method that allows us to keep track of the round-off error and does not require previous knowledge of the exact solution. As an application we provide a rigorous proof of the construction/existence of…
Passive investing can incur hidden costs due to market timing inefficiencies.
Gradient descent benefits from tangent kernel advantages under specific conditions.
In this report, we derive a non-negative series expansion for the Jensen-Shannon divergence (JSD) between two probability distributions. This series expansion is shown to be useful for numerical calculations of the JSD, when the probability distributions are nearly equal, and for which, consequently, small numerical er…
Binary classification improves with a small fraction of corrupted labels.
Random feature model shows slow self-correction of generalization gap.
We study the effects of approximate inference on the performance of Thompson sampling in the -armed bandit problems. Thompson sampling is a successful algorithm for online decision-making but requires posterior inference, which often must be approximated in practice. We show that even small constant inference error …
Fast algorithm recovers principal eigenvector from noisy matrices.
The CLT fails for LLM evaluations with small data, leading to underestimation of uncertainty.
Paper proposes efficient communication scheme for statistical learning.
This research evaluates generalization measures in deep learning.
For binary classification we establish learning rates up to the order of for support vector machines (SVMs) with hinge loss and Gaussian RBF kernels. These rates are in terms of two assumptions on the considered distributions: Tsybakov's noise assumption to establish a small estimation error, and a new geometr…
Study evaluates uncertainty quantification for atomistic neural networks, revealing complex relationships between error and uncertainty.
Dimensionality reduction is a first step of many machine learning pipelines. Two popular approaches are principal component analysis, which projects onto a small number of well chosen but non-interpretable directions, and feature selection, which selects a small number of the original features. Feature selection can be…
Study on Transfer Elastic Net error bounds and grouping effect.
Small initialization improves tensor recovery from noisy data.
We propose to study the generalization error of a learned predictor in terms of that of a surrogate (potentially randomized) predictor that is coupled to and designed to trade empirical risk for control of generalization error. In the case where interpolates the data, it is interesting to con…
Identifies bilinear systems from a single trajectory with optimal sample complexity.
Neural networks can approximate rectifiable measures with small error.
Paper analyzes Gibbs and Langevin Monte Carlo for interpolation regime, showing generalization from low errors.
We propose rectified factor networks (RFNs) to efficiently construct very sparse, non-linear, high-dimensional representations of the input. RFN models identify rare and small events in the input, have a low interference between code units, have a small reconstruction error, and explain the data covariance structure. R…
Gradient descent fails to learn simple neural networks efficiently.
We design a new algorithm for the Euclidean -means problem that operates in the local model of differential privacy. Unlike in the non-private literature, differentially private algorithms for the -means objective incur both additive and multiplicative errors. Our algorithm significantly reduces the additive erro…
Study proposes a stopping criterion for active learning based on error stability.
Stochastic gradient descent updates parameters with summation gradient computed from a random data batch. This summation will lead to unbalanced training process if the data we obtained is unbalanced. To address this issue, this paper takes the error variance and error mean both into consideration. The adaptively adjus…
This work presents a technique for statistically modeling errors introduced by reduced-order models. The method employs Gaussian-process regression to construct a mapping from a small number of computationally inexpensive `error indicators' to a distribution over the true error. The variance of this distribution can be…
Suppose some classifiers are selected from a set of hypothesis classifiers to form an equally-weighted ensemble that selects a member classifier at random for each input example. Then the ensemble has an error bound consisting of the average error bound for the member classifiers, a term for selectivity that varies fro…
Ex ante forecast outcomes should be interpreted as counterfactuals (potential histories), with errors as the spread between outcomes. Reapplying measurements of uncertainty about the estimation errors of the estimation errors of an estimation leads to branching counterfactuals. Such recursions of epistemic uncertainty …
Hardness proof for agnostically learning halfspaces from worst-case lattice problems.
In this paper, we study the problem of sparse multiple kernel learning (MKL), where the goal is to efficiently learn a combination of a fixed small number of kernels from a large pool that could lead to a kernel classifier with a small prediction error. We develop an efficient algorithm based on the greedy coordinate d…
Geometric framework links clustering accuracy to structural recovery.
A new data-adaptive prior stabilizes kernel learning in operators.
The CUR matrix decomposition is an important extension of Nyström approximation to a general matrix. It approximates any data matrix in terms of a small number of its columns and rows. In this paper we propose a novel randomized CUR algorithm with an expected relative-error bound. The proposed algorithm has the advanta…