A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Optimal regularization can prevent the double descent phenomenon in learning models.
problem The double descent phenomenon in learning models, where test performance is non-monotonic in sample size and model size.
method Theoretical and empirical study of optimal ℓ2 regularization for linear regression models and neural networks.
result Optimally-tuned ℓ2 regularization achieves monotonic test performance for certain models and mitigates the double descent phenomenon for more general models.
We show that a variety of modern deep learning tasks exhibit a "double-descent" phenomenon where, as we increase model size, performance first gets worse and then gets better. Moreover, we show that double descent occurs not just as a function of model size, but also as a function of the number of training epochs. We u…
Double descent refers to the phase transition that is exhibited by the generalization error of unregularized learning models when varying the ratio between the number of parameters and the number of training samples. The recent success of highly over-parameterized machine learning models such as deep neural networks ha…
Study on double descent behavior in two-layer neural networks for binary classification.
problem Understanding the double descent phenomenon in model test error.
method Two-layer neural network with ReLU activation for binary classification. Quantified model size by sample-to-dimension ratio. Empirical risk minimization using Convex Gaussian Min Max Theorem.
result Observed and investigated the double descent behavior of model test error.
Reinforcement learning studies how to balance exploration and exploitation in real-world systems, optimizing interactions with the world while simultaneously learning how the world operates. One general class of algorithms for such learning is the multi-armed bandit setting. Randomized probability matching, based upon …
The "double descent" risk curve was proposed to qualitatively describe the out-of-sample prediction accuracy of variably-parameterized machine learning models. This article provides a precise mathematical analysis for the shape of this curve in two simple data models with the least squares/least norm predictor. Specifi…
Double descent in portfolio optimization shows improved performance with complexity, then declines, due to overfitting.
problem Improving portfolio optimization performance with model complexity.
method Investigates the relationship between model complexity and out-of-sample performance in mean-variance portfolio optimization.
result Performance of low-dimensional models initially improves with complexity but declines due to overfitting. High-dimensional models show double ascent Sharpe ratio curve.
In this paper, we propose a Double Thompson Sampling (D-TS) algorithm for dueling bandit problems. As indicated by its name, D-TS selects both the first and the second candidates according to Thompson Sampling. Specifically, D-TS maintains a posterior distribution for the preference matrix, and chooses the pair of arms…
We address the problem of multi-class classification in the case where the number of classes is very large. We propose a double sampling strategy on top of a multi-class to binary reduction strategy, which transforms the original multi-class problem into a binary classification problem over pairs of examples. The aim o…
Sparse coding is a crucial subroutine in algorithms for various signal processing, deep learning, and other machine learning applications. The central goal is to learn an overcomplete dictionary that can sparsely represent a given input dataset. However, a key challenge is that storage, transmission, and processing of …
The aim of this paper is to establish two fundamental measure-metric properties of particular random geometric graphs. We consider ε-neighborhood graphs whose vertices are drawn independently and identically distributed from a common distribution defined on a regular submanifold of RK. We show t…
New findings challenge the traditional U-shaped curve of model complexity and error, revealing a second descent in error as model size increases.
problem The traditional U-shaped curve of model complexity and prediction error is incomplete, with recent work suggesting a second descent in error as model size increases.
method Careful consideration of multiple complexity axes and a nonparametric statistics perspective were used to interpret the observed double descent curves.
result The observed double descent curves in classical statistical machine learning methods fold back into traditional convex shapes, resolving tensions with statistical intuition.
A Klein surface is a surface with a dianalytic structure. A double of a Klein surface X is a Klein surface Y such that there is a degree two morphism (of Klein surfaces) Y→X. There are many doubles of a given Klein surface and among them the so-called natural doubles which are: the complex double, the …
Develops a test for conditional local independence of counting processes.
problem Testing the hypothesis of conditional local independence among continuous time stochastic processes.
method Introduces a new functional parameter, the Local Covariance Measure (LCM), and proposes a test called (X)-LCT using nonparametric estimators and sample splitting or cross-fitting.
result The (X)-LCT test can be controlled uniformly with modest rates, and it works well without restrictive parametric assumptions.
We define a general notion of abstract double Lie algebroid. We show (1) that the double Lie algebroid of a double Lie groupoid is a double Lie algebroid in this sense; (2) that the double cotangent constructed from Lie algebroid structures on a vector bundle A and its dual A* is a double Lie algebroid if and only if (…
The word `double' was used by Ehresmann to mean `an object X in the category of all X'. Double categories, double groupoids and double vector bundles are instances, but the notion of Lie algebroid cannot readily be doubled in the Ehresmann sense, since a Lie algebroid bracket cannot be defined diagrammatically. In this…
We propose a generalized double Pareto prior for Bayesian shrinkage estimation and inferences in linear models. The prior can be obtained via a scale mixture of Laplace or normal distributions, forming a bridge between the Laplace and Normal-Jeffreys' priors. While it has a spike at zero like the Laplace density, it al…
We define double principal bundles (DPBs), for which the frame bundle of a double vector bundle, double Lie groups and double homogeneous spaces are basic examples. It is shown that a double vector bundle can be realized as the associated bundle of its frame bundle. Also dual structures, gauge transformations and conne…