This work improves neural network calibration using explicit regularization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AIR-Net adapts low-rank regularization dynamically for better image completion.
Combining explicit and implicit regularization improves deep learning performance without needing depth.
Noise injection before gradient steps helps in regularization for neural networks.
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an -path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regula…
Dropout introduces both explicit and implicit regularization effects.
We prove that pseudo-holomorphic discs attached to a maximal totally real submanifold inherit their regularity from the regularity of the submanifold and of the almost complex structure. The proof is based on the computation of an explicit lower bound for the Kobayashi metric in almost complex manifolds, which also yie…
Study shows how networks converge to minimum norm solutions with regularization.
MARL algorithm uses regularization to avoid explicit structures, improving performance.
A new method bridges explicit and implicit deep generative models using Stein discrepancy.
The purpose of this paper is to give explicit methods for bounding the number of vertices of finite -regular graphs with given second eigenvalue. Let be a finite -regular graph and the second largest eigenvalue of its adjacency matrix. It follows from the well-known Alon-Boppana Theorem, that for any…
Ahern and Rudin have given an explicit construction of a totally real embedding of in . As a generalization of their example, we give an explicit example of a CR regular embedding of in . Consequently, we show that the odd dimensional sphere with admits…
Gradient descent implicitly regularizes neural networks by penalizing large loss gradients.
New method for training deep neural networks with regularization, converging to better generalization.
Lower discount factors act as a regularizer in RL, improving performance.
No regularization needed for InLDL, achieving efficient and effective model.
Many statistical estimators for high-dimensional linear regression are M-estimators, formed through minimizing a data-dependent square loss function plus a regularizer. This work considers a new class of estimators implicitly defined through a discretized gradient dynamic system under overparameterization. We show that…
We introduce -regular maps, which generalize two previously studied classes of maps: affinely -regular maps and totally skew embeddings. We exhibit some explicit examples and obtain bounds on the least dimension of a Euclidean space into which a manifold can be embedded by a -regular map. The problem c…
This paper presents an asynchronous incremental aggregated gradient algorithm and its implementation in a parameter server framework for solving regularized optimization problems. The algorithm can handle both general convex (possibly non-smooth) regularizers and general convex constraints. When the empirical data loss…
Efficient echo state network with explicit memory performs well on benchmark tasks.
Research provides explicit NPV expressions for double barrier strategies.
The paper studies convergence rates of Tsallis entropic regularization in optimal transport.
Choquet regularization improves exploration in RL.
We study the set of critical exponents of discrete groups acting on regular trees. We prove that for every real number between and , there is a discrete subgroup acting without inversion on a -regular tree whose critical exponent is equal to . Explicit construction of edge-index…
Kernel ridgeless regression with random features shows good generalization without explicit regularization.
Study pseudo-laplacians and ζ(1) for spinor bundles over Riemann surfaces.
Let F be a closed orientable surface. We give an explicit formula for the number mod 2 of quadruple points occurring in any generic regular homotopy between any two regularly homotopic embeddings e,e':F -> R^3. The formula is in terms of homological data extracted from the two embeddings.
We prove an explicit and sharp upper bound for the Castelnuovo-Mumford regularity of an FI-module V in terms of the degrees of its generators and relations. We use this to refine a result of Putman on the stability of homology of congruence subgroups, extending his theorem to previously excluded small characteristics a…
New approach stabilizes GANs by leveraging implicit competitive regularization.
Evolutoids of surfaces defined as line envelopes, studied using singularity theory.
A continuous map from R^m to R^N or from C^m to C^N is called k-regular if the images of any points are linearly independent. Given integers m and k a problem going back to Chebyshev and Borsuk is to determine the minimal value of N for which such maps exist. The methods of algebraic topology provide lower bounds f…
Several works have aimed to explain why overparameterized neural networks generalize well when trained by Stochastic Gradient Descent (SGD). The consensus explanation that has emerged credits the randomized nature of SGD for the bias of the training process towards low-complexity models and, thus, for implicit regulari…
We establish sharp regularity and Fredholm theorems for the \bar{\partial}_b-Neumann problem on domains satisfying some non-generic geometric conditions. We use these domains to construct explicit examples of bad behaviour of the Kohn Laplacian: it is not always hypoelliptic up to the boundary, its partial inverse is n…
Learning to approximate a separable function is hard, requiring many samples even with sparse networks.
The study examines the regularity of branched immersions using special coordinate systems.
Recent years have seen a flurry of activities in designing provably efficient nonconvex procedures for solving statistical estimation problems. Due to the highly nonconvex nature of the empirical loss, state-of-the-art procedures often require proper regularization (e.g. trimming, regularized cost, projection) in order…
The paper discusses how to improve machine learning models using partial differential equations.
Proposes a new model for image restoration combining deep learning and total variation.
Local regularization fails in transductive learning for some multiclass problems.
In this work we establish the equivalence of algorithmic regularization and explicit convex penalization for generic convex losses. We introduce a geometric condition for the optimization path of a convex function, and show that if such a condition is satisfied, the optimization path of an iterative algorithm on the un…
Just as an explicit parameterisation of system dynamics by state, i.e., a choice of coordinates, can impede the identification of general structure, so it is too with an explicit parameterisation of system dynamics by control. However, such explicit and fixed parameterisation by control is commonplace in control theory…
Study the Lax equation in infinite-dimensional Lie algebras and Lie groups.
Gradient descent recovers principal components of overparametrized asymmetric matrices without explicit regularization.
New iterative regularization method tackles non-smooth, non-strongly convex functionals.
Batch Normalization (BN) improves both convergence and generalization in training neural networks. This work understands these phenomena theoretically. We analyze BN by using a basic block of neural networks, consisting of a kernel layer, a BN layer, and a nonlinear activation function. This basic network helps us unde…
We define an infinite series of translation coverings of Veech's double-n-gon for odd n greater or equal to 5 which share the same Veech group. Additionally we give an infinite series of translation coverings with constant Veech group of a regular n-gon for even n greater or equal to 8. These families give rise to expl…
Paper develops a new probabilistic method for American options using entropy regularization.
Introduces self-regularization for analyzing learning algorithms.