Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

60120179239 · Jun 202019922001200920172026
48 results for double layer

This paper explains double descent in linear neural networks, identifying new factors.

problem Understanding double descent in linear neural networks.
method Gradient flow derivation and necessary conditions for double descent.
result Singular values of input-output covariance matrix are important for double descent in two-layer models.

The paper constructs triangulations for double twist knots using geometric methods.

problem Constructing explicit triangulations of double twist knots.
method Using triangulating Dehn fillings, layered solid tori, and their double covers.
result Proves both triangulations are geometric, using conjecturally minimal triangulation to present A-polynomial equations.

Double descent risk in L2-regularized models explained and mitigated.

problem Risk of overparameterized models in machine learning.
method Analysis of L2-regularized models, two-layer neural networks, and CNNs.
result Double descent risk in L2-regularized models can be explained and mitigated by adjusting regularization strengths.

Study on double descent behavior in two-layer neural networks for binary classification.

problem Understanding the double descent phenomenon in model test error.
method Two-layer neural network with ReLU activation for binary classification. Quantified model size by sample-to-dimension ratio. Empirical risk minimization using Convex Gaussian Min Max Theorem.
result Observed and investigated the double descent behavior of model test error.

Improves early stopping in deep networks by adjusting stepsizes.

problem Epoch-wise double descent in deep networks.
method Analytical and empirical study of bias-variance tradeoffs in different network layers.
result Eliminating epoch-wise double descent through adjusting stepsizes of different layers improves early stopping performance.

Proves invertibility of layer potentials for generalized Stokes operators on smooth domains.

problem Invertibility of layer potentials for generalized Stokes operators on smooth domains.
method Developed algebra toolkit to handle layer operators' limit and jump relations; proved Fredholm property and invertibility.
result Proves invertibility of layer potentials for generalized Stokes operators on smooth domains.

Deep learning models can overfit noisy data without losing generalization.

problem Understanding the generalization of deep learning models in noisy data.
method Empirical investigation of epoch-wise double descent in fully connected neural networks trained on CIFAR-10 with 30% label noise.
result The model achieves strong re-generalization on test data after overfitting noisy training data, corresponding to a 'benign overfitting' state.

Study compares random and learned features in deep Bayesian linear models.

problem Understanding how feature learning affects generalization in deep learning.
method Comparing deep random feature models to deep networks with trained layers.
result Random feature models can display double-descent behavior, while deep networks do not.

Annealing Double-Head calibrates deep neural networks during training.

problem Overestimation or underestimation of predictive confidence in deep neural networks.
method An additional calibration head and Annealing technique to dynamically scale logits.
result State-of-the-art model calibration performance achieved without post-processing.

This paper explains why double descent sometimes occurs weakly or not at all from an optimization perspective.

problem Understanding the role of optimization in the phenomenon of double descent.
method Investigates model-wise double descent from an optimization perspective, proposing a unified explanation for its occurrence.
result Model-wise double descent is observed if and only if the optimizer can find a sufficiently low-loss minimum.

The paper explores how noise in features can lead to benign overfitting in machine learning models.

problem Understanding the conditions for benign overfitting in machine learning models.
method Examined random feature models, specifically two-layer neural networks with fixed first layer weights, and analyzed the role of noise in features.
result Noise in features plays an important implicit regularization role in the phenomenon of benign overfitting.

Despite their prevalence in neural networks we still lack a thorough theoretical characterization of ReLU layers. This paper aims to further our understanding of ReLU layers by studying how the activation function ReLU interacts with the linear component of the layer and what role this interaction plays in the success …

2018-12-06abs ↗pdf ↗

Random Transformers behave like polynomial models in ICL with asymptotic growth.

problem Understanding in-context learning capabilities of pretrained Transformers.
method Asymptotic analysis of a random Transformer with a fixed first layer and a trained second layer, considering growth in context length, input dimension, hidden dimension, and training parameters.
result The random Transformer's ICL error is equivalent to a finite-degree Hermite polynomial model.

Deep networks preferentially learn shared features, avoiding memorization in early layers.

problem Understanding how deep neural networks generalize vs. memorize training data.
method Replica-based mean field geometric analysis of deep neural networks.
result Deep layers predominantly memorize, while early layers are minimally affected.

Analyzes generalization error in generalized linear models, explaining double descent phenomenon.

problem Understanding generalization of machine learning models in high dimensions.
method Develops a framework to characterize asymptotic generalization error for generalized linear models.
result Rigorously explains the double descent phenomenon in generalized linear models.

We investigate the equation (ΔHn)γw=f(w)inHn,(-Δ_{\mathbb H^n})^γ w=f(w)\quad in \mathbb H^{n}, where (ΔHn)γ(-Δ_{\mathbb H^n})^γ corresponds to the fractional Laplacian on hyperbolic space for γ(0,1)γ\in (0,1) and ff is a smooth nonlinearity that typically comes from a double well potential. We prove the existence of heteroclinic connecti…

2012-12-31abs ↗pdf ↗

This study shows ESG ratings reduce equity crash risk during market downturns.

problem Decoupling of alpha from tail risk resilience in traditional models.
method Double Machine Learning for structural deconfounding, state-dependent analysis.
result High ESG ratings reduce crash incidence during systemic drawdowns.

Improved defect detection in layered materials using signal separation methods.

problem Challenging defect detection due to strong clutter in layered structures.
method Joint rank and sparsity minimization with an iteratively reweighted nuclear and 1\ell_1-norm approach, combined with deep learning for parameter optimization.
result The proposed approach outperforms conventional methods in terms of accuracy and speed of convergence.

Neural networks exhibit unimodal variance with model complexity, improving generalization.

problem The classical bias-variance trade-off does not apply to neural networks, leading to better generalization with larger models.
method Measured bias and variance of neural networks, confirmed empirically and theoretically.
result Neural networks show unimodal variance, leading to a double descent risk curve.

Overparametrized models are vulnerable to adversarial perturbations, affecting robust generalization.

problem Understanding how overparametrization impacts robustness in adversarial training.
method Analyzing random features regression models with a precise asymptotic formula.
result High overparametrization can hurt robust generalization in adversarially trained models.

New theory explains how overparametrized neural networks generalize well without bias-variance trade-off.

problem Overparametrized neural networks generalize well despite classical bias-variance trade-off.
method Nonasymptotic generalization theory for two-layer neural networks with ReLU activation, incorporating scaled variation regularization.
result Prediction bounds for all network widths reproduce the double descent phenomenon, and overparametrized models are nearly minimax optimal.

We study minimal energy problems for strongly singular Riesz kernels on a manifold. Based on the spatial energy of harmonic double layer potentials, we are motivated to formulate the natural regularization of such problems by switching to Hadamard's partie finie integral operator which defines a strongly elliptic pseud…

2016-02-27abs ↗pdf ↗

Approximately, 50 million people in the world are affected by epilepsy. For patients, the anti-epileptic drugs are not always useful and these drugs may have undesired side effects on a patient's health. If the seizure is predicted the patients will have enough time to take preventive measures. The purpose of this work…

2019-12-13abs ↗pdf ↗

A Klein surface is a surface with a dianalytic structure. A double of a Klein surface XX is a Klein surface YY such that there is a degree two morphism (of Klein surfaces) YXY\rightarrow X. There are many doubles of a given Klein surface and among them the so-called natural doubles which are: the complex double, the …

2014-04-03abs ↗pdf ↗

Future video prediction is an ill-posed Computer Vision problem that recently received much attention. Its main challenges are the high variability in video content, the propagation of errors through time, and the non-specificity of the future frames: given a sequence of past frames there is a continuous distribution o…

2017-12-01abs ↗pdf ↗

The paper uses machine learning to optimize rework policies in semiconductor manufacturing.

problem Optimizing rework steps to increase yield without increasing costs.
method Applied double/debiased machine learning (DML) to estimate treatment effects.
result Derived optimal rework policies and estimated their value empirically.

We define a general notion of abstract double Lie algebroid. We show (1) that the double Lie algebroid of a double Lie groupoid is a double Lie algebroid in this sense; (2) that the double cotangent constructed from Lie algebroid structures on a vector bundle A and its dual A* is a double Lie algebroid if and only if (…

1998-08-17abs ↗pdf ↗

Derives asymptotic generalization error for large-margin classifiers.

problem Understanding the generalization error of large-margin classifiers.
method Statistical physics replica method for deriving asymptotic expression.
result Establishes phase transition boundary for class separability.

We study reinforcement learning (RL) in high dimensional episodic Markov decision processes (MDP). We consider value-based RL when the optimal Q-value is a linear function of d-dimensional state-action feature representation. For instance, in deep-Q networks (DQN), the Q-value is a linear function of the feature repres…

2018-02-13abs ↗pdf ↗

New insights into how overfitting affects neural networks' performance.

problem Understanding the generalization of overfitted two-layer neural networks.
method Analyzing the NTK model with ReLU activation, focusing on min 2\ell_2-norm solutions.
result Generalization error of overfitted NTK models approaches a small limiting value, even with infinite neurons and samples.

Study shows label noise impacts neural representations' information content, revealing double descent behavior.

problem Impact of label noise on neural network hidden representations.
method Information Imbalance proxy of conditional mutual information to compare hidden representations.
result Representations learned with noisy labels are more informative than those with clean labels in the underparameterized regime, and equally informative in the overparameterized regime.

We define double principal bundles (DPBs), for which the frame bundle of a double vector bundle, double Lie groups and double homogeneous spaces are basic examples. It is shown that a double vector bundle can be realized as the associated bundle of its frame bundle. Also dual structures, gauge transformations and conne…

2016-11-02abs ↗pdf ↗

This paper establishes an equivalence between transitive double Lie algebroids and core diagrams.

problem Understanding and characterizing transitive double Lie algebroids.
method Using core diagrams and equivalence of transitive core diagrams with transitive double Lie groupoids.
result Transitive double Lie algebroids are completely determined by their core diagrams.

Kernel methods and MLPs perform similarly to linear models in high dimensions.

problem Understanding the performance of kernel methods and MLPs in high-dimensional settings.
method Analysis of kernel methods and MLPs in a high-dimensional regime with proportional asymptotics.
result Linear models are optimal in high-dimensional settings when data is generated by kernel models with nonlinear relationships.

We define an abstract notion of double Lie algebroid, which includes as particular cases: (1) the double Lie algebroid of a double Lie groupoid in the sense of the author, such as the iterated tangent bundle of an ordinary manifold, and various iterated tangent/cotangent constructions in symplectic and Poisson geometry…

2000-11-24abs ↗pdf ↗

We develop new algebraic methods refining the Witt group of linking forms and Ranicki's torsion algebraic L-groups into double Witt groups and double L-groups. At each prime ideal of the underlying ring, our double Witt groups capture infinitely many more integral signatures of the linking form than the single Witt gro…

2015-03-24abs ↗pdf ↗

A theory of double affine and special double affine bundles, i.e. differential manifolds with two compatible (special) affine bundle structures, is developed as an affine counterpart of the theory of double vector bundles. The motivation and basic examples come from Analytical Mechanics, where double affine bundles hav…

2009-04-14abs ↗pdf ↗