Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

79158237316 · Jun 202019922001200920172026
48 results for 2-regular points

New algorithmic view of ℓ2 regularization using ODEs and path-following methods.

problem Optimizing convex loss functions with ℓ2 regularization.
method Established an equivalence between ℓ2-regularized solution paths and ODEs, proposing path-following algorithms based on homotopy methods and numerical ODE solvers.
result The solution path can be viewed as a hybrid of gradient descent and Newton method, providing novel schemes to choose grid points and reducing computational cost.

AdamW optimizes a constrained loss with \ell_\infty norm constraint.

problem Understanding the optimization behavior of AdamW with \ell_\infty norm constraint.
method Analyzing AdamW as a smoothed version of SignGD and connecting it to Frank-Wolfe optimization.
result AdamW implicitly performs constrained optimization with \ell_\infty norm constraint.

Dropout is a simple but effective technique for learning in neural networks and other settings. A sound theoretical understanding of dropout is needed to determine when dropout should be applied and how to use it most effectively. In this paper we continue the exploration of dropout as a regularizer pioneered by Wager,…

2014-12-15abs ↗pdf ↗

Study L2L_2 regularization in deep networks, uncovering performance relations and proposing a training schedule.

problem Understanding and optimizing L2L_2 regularization in deep learning models.
method Empirical observations and theoretical analysis of gradient flow dynamics in infinitely wide networks.
result Empirical relations between model performance, L2L_2 coefficient, learning rate, and training steps; optimal regularization parameter prediction; improved training schedule.

This work proves L2L_2-regularized ERM controls smCE without post-hoc correction.

problem Calibration of predicted probabilities in machine learning models.
method Canonical L2L_2-regularized empirical risk minimization.
result Theoretical proof that smCE is controlled by ERM without post-hoc correction.

The paper introduces a novel method for training neural network Stein critics with staged L2L^2-regularization.

problem Learning to differentiate model distributions from observed data in high-dimensional settings.
method Developed a novel staging procedure for L2L^2 regularization over training time, leveraging the advantages of highly-regularized training at early times.
result Theoretical guarantees and empirical validation show that the method improves the approximation of the training dynamic by the kernel optimization, leading to faster convergence and better performance.

Safe screening rules reduce computation time in logistic regression with 02\ell_0-\ell_2 regularization.

problem Efficiently solving logistic regression with many features and regularization.
method Screening rules based on Fenchel dual lower bounds of strong conic relaxations.
result A high percentage of features can be safely removed before solving, leading to substantial speed-up.

A new method for joint eQTL mapping and gene network estimation.

problem Discovering SNP-gene relationships and gene-gene relationships in gene expression regulation.
method L1-2 regularized multi-task graphical lasso (L1-2 GLasso).
result Competitive performance on capturing true sparse structures of eQTL mapping and gene network.

Overparametrized neural networks can generalize well with proper regularization.

problem Generalization guarantee for noisy data in overparametrized neural networks.
method Nonparametric analysis of 2\ell_2-regularized GD trajectories.
result Achieving minimax optimal rate of L2L_2 estimation error with 2\ell_2 regularization.

SGD can jump from high rank minima to low rank minima in DLNs, but not back.

problem SGD's tendency to get stuck in high rank minima in DLNs.
method Analysis of the L2L_{2}-regularized loss function of DLNs and the definition of absorbing sets.
result SGD has a non-zero probability to jump from high rank minima to low rank minima but zero probability to jump back.

Proposes new attribution methods for trees with regularization.

problem Feature attribution for trees trained with regularization.
method Prediction Decomposition Attribution (PreDecomp) and TreeInner.
result TreeInner shows state-of-the-art feature selection performance.

In this paper we study constant scalar curvature equation (CSCK), a nonlinear fourth order elliptic equation, and its weak solutions on Kähler manifolds. We first define a notion of weak solution of CSCK for an LL^\infty Kähler metric. The main result is to show that such a weak solution (with uniform LL^\infty bound…

2017-05-03abs ↗pdf ↗

Autoencoders are a deep learning model for representation learning. When trained to minimize the distance between the data and its reconstruction, linear autoencoders (LAEs) learn the subspace spanned by the top principal directions but cannot learn the principal directions themselves. In this paper, we prove that $L_2…

2019-01-23abs ↗pdf ↗

Improving generalization is one of the main challenges for training deep neural networks on classification tasks. In particular, a number of techniques have been proposed, aiming to boost the performance on unseen data: from standard data augmentation techniques to the 2\ell_2 regularization, dropout, batch normalizat…

2019-07-19abs ↗pdf ↗

Statistical analysis of regularization in continual learning tasks.

problem Understanding how regularization affects model performance in sequential learning.
method Derivation of convergence rates, iterative update formula, and optimal hyperparameters for generalized ℓ2-regularization.
result Optimal hyperparameters balance forward and backward knowledge transfer, improving model performance.

Weight decay is one of the standard tricks in the neural network toolbox, but the reasons for its regularization effect are poorly understood, and recent results have cast doubt on the traditional interpretation in terms of L2L_2 regularization. Literal weight decay has been shown to outperform L2L_2 regularization for…

2018-10-29abs ↗pdf ↗

Regularization can improve machine learning models' robustness against poisoning attacks.

problem Poisoning attacks degrade machine learning models' performance by manipulating a fraction of the training data.
method Proposes a novel optimal attack formulation considering the effect of hyperparameters on regularization, leading to better evaluation of robustness.
result Demonstrates that L2L_2 regularization can help mitigate the impact of poisoning attacks.

We give an extensive treatment of the Constant Mean Curvature (CMC) Einstein flow from the point of view of the Bel-Robinson energies. The article, in particular, stresses on estimates showing how the Bel-Robinson energies and the volume of the evolving states control intrinsically the flow along evolution. The treatme…

2007-05-21abs ↗pdf ↗

The paper proves regularity for varifolds with bounded anisotropic mean curvature.

problem Regularity of varifolds with bounded anisotropic mean curvature.
method Local anisotropic regularity theorem and touching balls approach.
result Varifolds can be covered by countably many C2C^2-regular submanifolds.

Via Gauge theory, we give a new proof of partial regularity for harmonic maps in dimension m>2 into arbitrary targets. This proof avoids the use of adapted frames and permits to consider targets of "minimal" C^2 regularity. The proof we present moreover extends to a large class of elliptic systems of quadratic growth.

2006-04-28abs ↗pdf ↗

New MIP framework solves high-dimensional 02\ell_0\ell_2-regularized regression problems.

problem Exact computation of 02\ell_0\ell_2-regularized regression estimators is challenging for large pp.
method Specialized nonlinear branch-and-bound (BnB) framework with first-order optimization.
result Achieves speedups of at least 5000x compared to state-of-the-art exact methods.

The paper constructs new non-trivial harmonic maps into higher-dimensional target manifolds.

problem Existence of non-trivial harmonic maps into higher-dimensional target manifolds.
method Perturbative argument, refined neck-analysis, energy identity, min-max problems.
result Construction of an infinite family of new null-homotopic nn-harmonic nn-spheres.

Optimal regularization can prevent the double descent phenomenon in learning models.

problem The double descent phenomenon in learning models, where test performance is non-monotonic in sample size and model size.
method Theoretical and empirical study of optimal 2\ell_2 regularization for linear regression models and neural networks.
result Optimally-tuned 2\ell_2 regularization achieves monotonic test performance for certain models and mitigates the double descent phenomenon for more general models.

We study metric spaces homeomorphic to the 2-sphere, and find conditions under which they are quasisymmetrically homeomorphic to the standard 2-sphere. As an application of our main theorem we show that an Ahlfors 2-regular, linearly locally contractible metric 2-sphere is quasisymmetrically homeomorphic to the standar…

2001-07-24abs ↗pdf ↗

Multiple generalized additive models (GAMs) are a type of distributional regression wherein parameters of probability distributions depend on predictors through smooth functions, with selection of the degree of smoothness via L2L_2 regularization. Multiple GAMs allow finer statistical inference by incorporating explana…

2018-09-25abs ↗pdf ↗

The paper analyzes SGD with dropout regularization in linear models, proving asymptotic properties and providing inference tools.

problem Analyzing the behavior of SGD with dropout regularization in linear models.
method Establishing geometric-moment contraction (GMC) and proving quenched central limit theorems (CLT).
result The existence of a unique stationary distribution and asymptotic normality results for SGD with dropout.

New method identifies network structure without regularization for sparse teacher couplings.

problem Identifying network structure in inverse Ising problems with model mismatch.
method Ridge linear regression with two-stage estimator.
result Perfect identification of network structure possible without regularization for sparse teacher couplings.

Early stopping improves logistic regression's calibration and consistency in high dimensions.

problem Improving the statistical performance of gradient descent in overparameterized logistic regression.
method Investigates the effects of early stopping on gradient descent in logistic regression.
result Early-stopped gradient descent is well-calibrated and statistically consistent, while asymptotic gradient descent is not.

Study shows how classifiers can approach Bayes error in high-dimensional settings.

problem Generalization error in high-dimensional perceptrons.
method Proved a formula for generalization error using convex optimization and observed that logistic and hinge regression can approach Bayes error closely.
result Logistic and hinge regression can approach Bayes-optimal generalization error closely in high-dimensional settings.

The paper studies how adding an ℓ2 penalty affects network embeddings.

problem The impact of ℓ2 regularization on network embeddings.
method Analyzes the asymptotic behavior of ℓ2 regularized node2vec embeddings under graphon theory.
result The learned embeddings asymptotically form a graphon with a nuclear-norm-type penalty.

We prove an analogue of Thurston's h-principle for 22-dimensional foliations on manifolds of dimension bigger or equal to 44, in the presence of a fiber-wise non-degenerate 22-form. This helps us understand the flexibility of rank 22 regular Poisson structures on open manifolds with dimension bigger or equal to 44

2016-11-29abs ↗pdf ↗