Gradient descent implicitly follows regularization for general losses.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New fairness approach removes direct effects of unprivileged groups through causal regularization.
CASTLE learns causal DAG to improve model generalization.
Proves approximation and interpolation for regular immersions directed by algebraically elliptic cones.
New method improves speed of estimating bivariate functional data.
Identifying the underlying directional relations from observational time series with nonlinear interactions and complex relational structures is key to a wide range of applications, yet remains a hard problem. In this work, we introduce a novel minimum predictive information regularization method to infer directional r…
New regularization method corrects over-shrinkage in small data regression.
Autoencoders are a deep learning model for representation learning. When trained to minimize the distance between the data and its reconstruction, linear autoencoders (LAEs) learn the subspace spanned by the top principal directions but cannot learn the principal directions themselves. In this paper, we prove that $L_2…
Community detection has been one of the central problems in network studies and directed network is particularly challenging due to asymmetry among its links. In this paper, we found that incorporating the direction of links reveals new perspectives on communities regarding to two different roles, source and terminal, …
We propose a new stochastic dual coordinate ascent technique that can be applied to a wide range of regularized learning problems. Our method is based on Alternating Direction Multiplier Method (ADMM) to deal with complex regularization functions such as structured regularizations. Although the original ADMM is a batch…
The paper develops methods to reduce deployment risk under dynamic covariate shifts.
Transfer learning have been frequently used to improve deep neural network training through incorporating weights of pre-trained networks as the starting-point of optimization for regularization. While deep transfer learning can usually boost the performance with better accuracy and faster convergence, transferring wei…
We show that directed minimal cones in (n+1)-dimensional Euclidean space which have at most one singularity are - besides the trivial cases: empty set, whole space - half spaces. Using blow-up techniques, this result can be used to get C^{1,lambda}-regularity for the measure-theoretic boundary of almost minimal Cacciop…
Regularization is a popular technique in machine learning for model estimation and avoiding overfitting. Prior studies have found that modern ordered regularization can be more effective in handling highly correlated, high-dimensional data than traditional regularization. The reason stems from the fact that the ordered…
Regularization plays a crucial role in supervised learning. Most existing methods enforce a global regularization in a structure agnostic manner. In this paper, we initiate a new direction and propose to enforce the structural simplicity of the classification boundary by regularizing over its topological complexity. In…
Regular variation provides a convenient theoretical framework to study large events. In the multivariate setting, the dependence structure of the positive extremes is characterized by a measure - the spectral measure - defined on the positive orthant of the unit sphere. This measure gathers information on the localizat…
PALMS reconstructs large-scale networks efficiently with parallel computing.
Bayesian -regularized least squares is a variable selection technique for high dimensional predictors. The challenge is optimizing a non-convex objective function via search over model space consisting of all possible predictor combinations. Spike-and-slab (a.k.a. Bernoulli-Gaussian) priors are the gold standard f…
Enhances deep learning by boosting generalization and convergence.
Sketchy reduces memory and compute requirements for adaptive regularization in deep learning.
Two definitions for the rectfiability of hypersurfaces in Heisenberg groups have been proposed: one based on -regular surfaces, and the other on Lipschitz images of subsets of codimension- vertical subgroups. The equivalence between these notions remains an open problem. Recent partial res…
New method prevents deep learning models from forgetting past tasks.
The support vector machine (SVM) was originally designed for binary classifications. A lot of effort has been put to generalize the binary SVM to multiclass SVM (MSVM) which are more complex problems. Initially, MSVMs were solved by considering their dual formulations which are quadratic programs and can be solved by s…
Paper studies heat flow for maps on manifolds, avoiding singularities.
The Finsleroid--Finsler space becomes regular when the norm of the input 1-form is taken to be an arbitrary positive scalar . By performing required direct evaluations, the respective spray coefficients have been obtained in a simple and transparent form. The adequate continuation into the regul…
Study of light function singularities on surfaces.
We study direct limits of Gelfand pairs of the form with nilpotent, in other words pairs for which is a commutative nilmanifold. First, we extend the criterion of \cite{W3} for a direct limit representation to be multiplicity free. Then w…
Proposes new stochastic algorithms for multi-objective optimization.
For a bounded domain of class , the properties are studied of fields of `good directions', that is the directions with respect to which can be locally represented as the graph of a continuous function. For any such domain there is a canonical smooth field of good direct…
A typical approach in estimating the learning rate of a regularized learning scheme is to bound the approximation error by the sum of the sampling error, the hypothesis error and the regularization error. Using a reproducing kernel space that satisfies the linear representer theorem brings the advantage of discarding t…
Regularization effect found in neural feature alignment.
Recent literature on online learning has focused on developing adaptive algorithms that take advantage of a regularity of the sequence of observations, yet retain worst-case performance guarantees. A complementary direction is to develop prediction methods that perform well against complex benchmarks. In this paper, we…
New method aggregates nodes in sparse graphical models.
Maximal regularity for nonuniformly parabolic problems with normal degeneration.
New autoencoder improves latent space learning by optimizing sliced Gromov-Wasserstein discrepancies.
GANs excel at learning high dimensional distributions, but they can update generator parameters in directions that do not correspond to the steepest descent direction of the objective. Prominent examples of problematic update directions include those used in both Goodfellow's original GAN and the WGAN-GP. To formally d…
A new PCR method using SVD with sparse regularization.
We use normal sections to relate the curvature locus of regular (resp. singular corank 1) 3-manifolds in (resp. ) with regular (resp. singular corank 1) surfaces in (resp. ). For example we show how to generate a Roman surface by a family of ellipses different to S…
We study the low-regularity (in-)extendibility of spacetimes within the synthetic-geometric framework of Lorentzian length spaces developed in [KS:17]. To this end, we introduce appropriate notions of geodesics and timelike geodesic completeness and prove a general inextendibility result. Our results shed new light on …
We study the existence and properties of metrics maximising the first Laplace eigenvalue among conformal metrics of unit volume on Riemannian surfaces. We describe a general approach to this problem and its higher eigenvalue versions via the direct method of calculus of variations. The principal results include the gen…
We prove existence of harmonic coordinates for the nonlinear Laplacian of a Finsler manifold and apply them in a proof of the Myers--Steenrod theorem for Finsler manifolds. Different from the Riemannian case, these coordinates are not suitable for studying optimal regularity of the fundamental tensor, nevertheless, we …
This paper deals with Prym eigenforms which are introduced previously by McMullen. We prove several results on the directional flow on those surfaces, related to complete periodicity (introduced by Calta). More precisely we show that any homological direction is algebraically periodic, and any direction of a regular cl…
Efficient inference for adaptive data with directional stability condition.
Adaptive regularization methods pre-multiply a descent direction by a preconditioning matrix. Due to the large number of parameters of machine learning problems, full-matrix preconditioning methods are prohibitively expensive. We show how to modify full-matrix adaptive regularization in order to make it practical and e…
This work improves neural network calibration using explicit regularization.
The use of convex regularizers allows for easy optimization, though they often produce biased estimation and inferior prediction performance. Recently, nonconvex regularizers have attracted a lot of attention and outperformed convex ones. However, the resultant optimization problem is much harder. In this paper, for a …
A new method solves sparse regularization problems efficiently and robustly.
Grokking occurs at numerical stability edge, requiring regularization to prevent.