2-regular points found in spaces with lower Ricci curvature bound.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithmic view of ℓ2 regularization using ODEs and path-following methods.
Establish C^{1,2} regularity of American value functions in Heston model
Given two Jordan curves in a Riemannian manifold, a minimal surface of annulus type bounded by these curves is described as the harmonic extension of a critical point of some functional (the Dirichlet integral) in a certain space of boundary parametrizations. The -regularity of the minimal surface of annulus t…
AdamW optimizes a constrained loss with norm constraint.
Dropout is a simple but effective technique for learning in neural networks and other settings. A sound theoretical understanding of dropout is needed to determine when dropout should be applied and how to use it most effectively. In this paper we continue the exploration of dropout as a regularizer pioneered by Wager,…
Study regularization in deep networks, uncovering performance relations and proposing a training schedule.
This work proves -regularized ERM controls smCE without post-hoc correction.
Paper solves Carathéodory's conjecture for -regular convex surfaces.
The paper introduces a novel method for training neural network Stein critics with staged -regularization.
Safe screening rules reduce computation time in logistic regression with regularization.
We show that any 2-valued C^{1, α} (α\in (0, 1)) function u = {u_{1}, u_{2}} on an open ball B in {\mathbb R}^{n} with values u_{1}, u_{2} \in {\mathbb R}^{k} whose graph, viewed as a varifold with multiplicity 2 at points where u_{1} = u_{2} and with multiplicity 1 at points where u_{1}, u_{2} are distinct, is station…
DNNs with regularization reveal feature learning dynamics and sparsity.
In this paper, we investigate a multivariate multi-response (MVMR) linear regression problem, which contains multiple linear regression models with differently distributed design matrices, and different regression and output vectors. The goal is to recover the support union of all regression vectors using -reg…
A new method for joint eQTL mapping and gene network estimation.
Optimal $C^{1,rac{1}{2}}$-regularity for -surfaces with free boundary.
Overparametrized neural networks can generalize well with proper regularization.
SGD can jump from high rank minima to low rank minima in DLNs, but not back.
Proposes new attribution methods for trees with regularization.
In this paper we study constant scalar curvature equation (CSCK), a nonlinear fourth order elliptic equation, and its weak solutions on Kähler manifolds. We first define a notion of weak solution of CSCK for an Kähler metric. The main result is to show that such a weak solution (with uniform bound…
Autoencoders are a deep learning model for representation learning. When trained to minimize the distance between the data and its reconstruction, linear autoencoders (LAEs) learn the subspace spanned by the top principal directions but cannot learn the principal directions themselves. In this paper, we prove that $L_2…
Improving generalization is one of the main challenges for training deep neural networks on classification tasks. In particular, a number of techniques have been proposed, aiming to boost the performance on unseen data: from standard data augmentation techniques to the regularization, dropout, batch normalizat…
Statistical analysis of regularization in continual learning tasks.
Weight decay is one of the standard tricks in the neural network toolbox, but the reasons for its regularization effect are poorly understood, and recent results have cast doubt on the traditional interpretation in terms of regularization. Literal weight decay has been shown to outperform regularization for…
By using the viewpoint of modern computational algebraic geometry, we explore properties of the optimization landscapes of the deep linear neural network models. After clarifying on the various definitions of "flat" minima, we show that the geometrically flat minima, which are merely artifacts of residual continuous sy…
The role of regularization, in the specific case of deep neural networks rather than more traditional machine learning models, is still not fully elucidated. We hypothesize that this complex interplay is due to the combination of overparameterization and high dimensional phenomena that take place during training …
Regularization can improve machine learning models' robustness against poisoning attacks.
We give an extensive treatment of the Constant Mean Curvature (CMC) Einstein flow from the point of view of the Bel-Robinson energies. The article, in particular, stresses on estimates showing how the Bel-Robinson energies and the volume of the evolving states control intrinsically the flow along evolution. The treatme…
The paper proves regularity for varifolds with bounded anisotropic mean curvature.
Regularization leads to balancedness in deep linear networks.
We give complete classification of C^2-regular and non-degenerate projectively Anosov flows on three dimensional manifolds. More precisely, we prove that such a flow on a connected manifold must be either an Anosov flow or represented as a finite union of -models.
Via Gauge theory, we give a new proof of partial regularity for harmonic maps in dimension m>2 into arbitrary targets. This proof avoids the use of adapted frames and permits to consider targets of "minimal" C^2 regularity. The proof we present moreover extends to a large class of elliptic systems of quadratic growth.
New MIP framework solves high-dimensional -regularized regression problems.
The optimization of a large random portfolio under the Expected Shortfall risk measure with an regularizer is carried out by analytical calculation. The regularizer reins in the large sample fluctuations and the concomitant divergent estimation error, and eliminates the phase transition where this error would …
The paper constructs new non-trivial harmonic maps into higher-dimensional target manifolds.
Optimal regularization can prevent the double descent phenomenon in learning models.
Regularization is a popular technique in machine learning for model estimation and avoiding overfitting. Prior studies have found that modern ordered regularization can be more effective in handling highly correlated, high-dimensional data than traditional regularization. The reason stems from the fact that the ordered…
We study metric spaces homeomorphic to the 2-sphere, and find conditions under which they are quasisymmetrically homeomorphic to the standard 2-sphere. As an application of our main theorem we show that an Ahlfors 2-regular, linearly locally contractible metric 2-sphere is quasisymmetrically homeomorphic to the standar…
Multiple generalized additive models (GAMs) are a type of distributional regression wherein parameters of probability distributions depend on predictors through smooth functions, with selection of the degree of smoothness via regularization. Multiple GAMs allow finer statistical inference by incorporating explana…
The paper analyzes SGD with dropout regularization in linear models, proving asymptotic properties and providing inference tools.
New method identifies network structure without regularization for sparse teacher couplings.
Early stopping improves logistic regression's calibration and consistency in high dimensions.
State-of-the-art subspace clustering methods are based on expressing each data point as a linear combination of other data points while regularizing the matrix of coefficients with , or nuclear norms. regularization is guaranteed to give a subspace-preserving affinity (i.e., there are no conne…
Sharp 3D Alexandrov inequality applied to volume-preserving flows.
We prove the following quantitative version of the celebrated Soap Bubble Theorem of Alexandrov. Let be a closed embedded hypersurface of , , and denote by the oscillation of its mean curvature. We prove that there exists a positive , depending on and upper …
Study shows how classifiers can approach Bayes error in high-dimensional settings.
The paper studies how adding an ℓ2 penalty affects network embeddings.
We prove an analogue of Thurston's h-principle for -dimensional foliations on manifolds of dimension bigger or equal to , in the presence of a fiber-wise non-degenerate -form. This helps us understand the flexibility of rank regular Poisson structures on open manifolds with dimension bigger or equal to …