This work improves neural network calibration using explicit regularization.
problem Improving predictive uncertainty in neural networks.
method Introducing a probabilistic calibration measure and exploring explicit regularization techniques.
result Explicit regularization improves log-likelihood and predictive uncertainty.
AIR-Net adapts low-rank regularization dynamically for better image completion.
problem Fixed low-rank regularization limits adaptability to different images.
method AIR-Net uses adaptive and implicit regularization parameterized by a dynamic Laplacian matrix.
result AIR-Net enhances implicit regularization and outperforms fixed methods in non-uniform missing data scenarios.
Combining explicit and implicit regularization improves deep learning performance without needing depth.
problem Improving deep learning performance without increasing model complexity.
method Proposes an explicit penalty to mirror implicit regularization bias in adaptive gradient optimizers.
result Single-layer networks can achieve low-rank approximations with similar performance to deep linear networks.
Noise injection before gradient steps helps in regularization for neural networks.
problem Improving generalization in overparametrized neural networks.
method Injecting small noise perturbations before computing gradient steps, especially in layer-wise fashion.
result Small noise perturbations can explicitly regularize neural networks without variance explosion.
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an ℓ2-path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regula…
Dropout introduces both explicit and implicit regularization effects.
problem Understanding the full impact of dropout regularization.
method Disentangled explicit and implicit regularization effects through experiments and analytic simplifications.
result Explicit and implicit regularization effects of dropout are distinct and can be characterized analytically.
We prove that pseudo-holomorphic discs attached to a maximal totally real submanifold inherit their regularity from the regularity of the submanifold and of the almost complex structure. The proof is based on the computation of an explicit lower bound for the Kobayashi metric in almost complex manifolds, which also yie…
Explicitly constructs CR regular embeddings of spheres in complex spaces.
problem Embedding spheres in complex spaces with CR regularity.
method Generalizes Ahern and Rudin's construction to higher dimensions.
result Odd dimensional spheres admit CR regular embeddings in complex spaces if and only if the dimension is even.
Study shows how networks converge to minimum norm solutions with regularization.
problem Interpolating between known regions in shallow ReLU networks.
method Investigates empirical risk minimizers and weight decay regularizers.
result Empirical risk minimizers converge to minimum norm interpolants under specific conditions.
MARL algorithm uses regularization to avoid explicit structures, improving performance.
problem Lack of effective reinforcement learning methods for multi-agent systems.
method MARQ uses regularization to promote structured exploration without explicit centralized structures.
result MARQ outperforms existing methods in multi-agent environments.
A new method bridges explicit and implicit deep generative models using Stein discrepancy.
problem Limitations of explicit and implicit deep generative models.
method Joint training framework that combines an explicit density estimator and an implicit sample generator via Stein discrepancy.
result The method improves the accuracy of density estimation and quality of generated samples.
The purpose of this paper is to give explicit methods for bounding the number of vertices of finite k-regular graphs with given second eigenvalue. Let X be a finite k-regular graph and μ1(X) the second largest eigenvalue of its adjacency matrix. It follows from the well-known Alon-Boppana Theorem, that for any…
Gradient descent implicitly regularizes neural networks by penalizing large loss gradients.
problem How to optimize deep neural networks without explicit regularization.
method Backward error analysis to calculate implicit gradient regularization and demonstrate its effectiveness empirically.
result Implicit gradient regularization biases gradient descent toward flat minima, improving model robustness and test errors.
New method for training deep neural networks with regularization, converging to better generalization.
problem Improving generalization of deep neural networks through explicit regularization.
method Regularizer Mirror Descent (RMD) method, inspired by convergence properties of stochastic mirror descent (SMD).
result RMD converges to a point close to the minimizer of the cost function, leading to better generalization performance.
Lower discount factors act as a regularizer in RL, improving performance.
problem Improving RL performance with limited data.
method Explicitly equating reduced discount factors to regularization terms.
result Regularization effectiveness depends on data properties.
No regularization needed for InLDL, achieving efficient and effective model.
problem InLDL struggles with performance degradation due to missing degrees.
method Proposes a model that uses label distribution as a prior, implicitly regularizing the learning process.
result Achieves competitive performance without explicit regularization.
Many statistical estimators for high-dimensional linear regression are M-estimators, formed through minimizing a data-dependent square loss function plus a regularizer. This work considers a new class of estimators implicitly defined through a discretized gradient dynamic system under overparameterization. We show that…
We introduce (k,l)-regular maps, which generalize two previously studied classes of maps: affinely k-regular maps and totally skew embeddings. We exhibit some explicit examples and obtain bounds on the least dimension of a Euclidean space into which a manifold can be embedded by a (k,l)-regular map. The problem c…
This paper presents an asynchronous incremental aggregated gradient algorithm and its implementation in a parameter server framework for solving regularized optimization problems. The algorithm can handle both general convex (possibly non-smooth) regularizers and general convex constraints. When the empirical data loss…
Efficient echo state network with explicit memory performs well on benchmark tasks.
problem Training differentiable neural computers is difficult and time-consuming.
method Echo state network with an explicit memory.
result Echo state network can recognize all regular languages, including those contractive networks cannot.
Research provides explicit NPV expressions for double barrier strategies.
problem Calculating expected NPVs of double barrier strategies for regular diffusions.
method Explicit expression using bivariate q-scale function with perturbation technique.
result Explicit expressions for expected NPVs are derived for certain cases.
The paper studies convergence rates of Tsallis entropic regularization in optimal transport.
problem Optimal transport with regularization.
method Γ-convergence and quantization/shadow arguments.
result Derives convergence rate of Tsallis entropic regularization.
Choquet regularization improves exploration in RL.
problem Improving exploration in reinforcement learning.
method Introducing Choquet regularizers to measure and manage exploration, reformulating RL problems and deriving explicit solutions.
result Explicit optimal distributions and Choquet regularizers for various exploratory samplers.
We study the set of critical exponents of discrete groups acting on regular trees. We prove that for every real number δ between 0 and 21logq, there is a discrete subgroup Γ acting without inversion on a (q+1)-regular tree whose critical exponent is equal to δ. Explicit construction of edge-index…
Kernel ridgeless regression with random features shows good generalization without explicit regularization.
problem Generalization of kernel ridgeless regression without explicit regularization.
method Investigation of ridgeless regression with random features and stochastic gradient descent, exploring the effect of random features error and spectral density optimization.
result Random features error exhibits the double-descent curve, leading to improved generalization.
Study pseudo-laplacians and ζ(1) for spinor bundles over Riemann surfaces.
problem Analyzing self-adjoint extensions of Dolbeault Laplacians on Riemann surfaces.
method Defined ζ-regularized determinants, introduced Robin mass, derived comparison formulas. result Explicit expressions for Robin mass in spinor bundles and scalar cases.
Let F be a closed orientable surface. We give an explicit formula for the number mod 2 of quadruple points occurring in any generic regular homotopy between any two regularly homotopic embeddings e,e':F -> R^3. The formula is in terms of homological data extracted from the two embeddings.
We prove an explicit and sharp upper bound for the Castelnuovo-Mumford regularity of an FI-module V in terms of the degrees of its generators and relations. We use this to refine a result of Putman on the stability of homology of congruence subgroups, extending his theorem to previously excluded small characteristics a…
New approach stabilizes GANs by leveraging implicit competitive regularization.
problem Stability issues in GAN training due to discriminator's exploitation of generator errors.
method Competitive Gradient Descent (CGD) for opponent-aware modeling of generator and discriminator.
result Significant improvement in GAN training stability without explicit regularization.
Evolutoids of surfaces defined as line envelopes, studied using singularity theory.
problem Defining and studying evolutoids of surfaces in 3D space.
method Explicit parametrization and singularity theory.
result Relations between surface geometry and its evolutoid.
A continuous map from R^m to R^N or from C^m to C^N is called k-regular if the images of any k points are linearly independent. Given integers m and k a problem going back to Chebyshev and Borsuk is to determine the minimal value of N for which such maps exist. The methods of algebraic topology provide lower bounds f…
Several works have aimed to explain why overparameterized neural networks generalize well when trained by Stochastic Gradient Descent (SGD). The consensus explanation that has emerged credits the randomized nature of SGD for the bias of the training process towards low-complexity models and, thus, for implicit regulari…
We establish sharp regularity and Fredholm theorems for the \bar{\partial}_b-Neumann problem on domains satisfying some non-generic geometric conditions. We use these domains to construct explicit examples of bad behaviour of the Kohn Laplacian: it is not always hypoelliptic up to the boundary, its partial inverse is n…
Learning to approximate a separable function is hard, requiring many samples even with sparse networks.
problem Learning the separable function x↦∑i=1dxi2 with limited samples. method Sparse neural networks vs. dense neural networks, explicit regularization.
result The sample complexity for dense networks is O(d2.5) with explicit regularization, better than O(d4). The study examines the regularity of branched immersions using special coordinate systems.
problem Understanding the regularity of branched immersions and their fundamental elements.
method Development and use of special coordinate systems to express maps with branch points, proving existence and regularity conditions for mean curvature vectors.
result Characterization and existence of special coordinate systems for branch immersions, proving regularity conditions for mean curvature vectors.
Recent years have seen a flurry of activities in designing provably efficient nonconvex procedures for solving statistical estimation problems. Due to the highly nonconvex nature of the empirical loss, state-of-the-art procedures often require proper regularization (e.g. trimming, regularized cost, projection) in order…
The paper discusses how to improve machine learning models using partial differential equations.
problem Improving the performance and generalization of machine learning models.
method The paper reframes implicit regularization techniques in deep learning as explicit gradient regularization using partial differential equations.
result Explicit regularization using PDEs can lead to better model performance and generalization.
Proposes a new model for image restoration combining deep learning and total variation.
problem Restoring images from limited data with low-rank constraints insufficient.
method Regularized Deep Matrix Factorized (RDMF) model using deep neural network's low-rank bias and total variation.
result Outperforms state-of-the-art models in image restoration from few observations.
Local regularization fails in transductive learning for some multiclass problems.
problem Whether local regularization can learn all transductive multiclass problems.
method Provided a negative answer by exhibiting a specific multiclass problem.
result Local regularization cannot learn all transductive multiclass problems.
In this work we establish the equivalence of algorithmic regularization and explicit convex penalization for generic convex losses. We introduce a geometric condition for the optimization path of a convex function, and show that if such a condition is satisfied, the optimization path of an iterative algorithm on the un…
Just as an explicit parameterisation of system dynamics by state, i.e., a choice of coordinates, can impede the identification of general structure, so it is too with an explicit parameterisation of system dynamics by control. However, such explicit and fixed parameterisation by control is commonplace in control theory…
Study the Lax equation in infinite-dimensional Lie algebras and Lie groups.
problem Investigate the Lax equation in infinite-dimensional Lie algebras and Lie groups.
method Derived integral expansions and generalized Baker-Campbell-Hausdorff formula for Lie groups.
result Explicit representation of product integral in terms of exponential map.
Gradient descent recovers principal components of overparametrized asymmetric matrices without explicit regularization.
problem Asymmetric matrix factorization under overparametrization with minimal rank assumptions.
method Vanilla gradient descent with small random initialization and proper early stopping.
result Gradient descent produces the best low-rank approximation without explicit regularization.
Batch Normalization (BN) improves both convergence and generalization in training neural networks. This work understands these phenomena theoretically. We analyze BN by using a basic block of neural networks, consisting of a kernel layer, a BN layer, and a nonlinear activation function. This basic network helps us unde…
New iterative regularization method tackles non-smooth, non-strongly convex functionals.
problem Tackles non-smooth, non-strongly convex functionals in regularization problems.
method Primal-dual algorithm with convergence and stability analysis.
result First iterative regularization procedure for non-smooth, non-strongly convex functionals.
We define an infinite series of translation coverings of Veech's double-n-gon for odd n greater or equal to 5 which share the same Veech group. Additionally we give an infinite series of translation coverings with constant Veech group of a regular n-gon for even n greater or equal to 8. These families give rise to expl…
Paper develops a new probabilistic method for American options using entropy regularization.
problem Finding optimal stopping times for American options with entropy regularization.
method Entropy-regularized penalization scheme based on Doob-Meyer-Mertens decomposition and reflected backward stochastic differential equations.
result Explicit convergence rates and policy improvement algorithm for American options.
Introduces self-regularization for analyzing learning algorithms.
problem Analyzing and optimizing learning algorithms without explicit regularization.
method Develops a self-regularization framework for learning algorithms.
result Provides statistical analysis and minmax-optimal rates for self-regularized algorithms.