In this work we study input gradient regularization of deep neural networks, and demonstrate that such regularization leads to generalization proofs and improved adversarial robustness. The proof of generalization does not overcome the curse of dimensionality, but it is independent of the number of layers in the networ…
This study connects Jacobian regularization to adversarial robustness and improves generalization.
problem Adversarial attacks make deep neural networks vulnerable.
method Developed a connection between Jacobian regularization and adversarial training, and established robust generalization gaps.
result Jacobian norms are related to both standard and robust generalization.
In this paper, we give a new generalization error bound of Multiple Kernel Learning (MKL) for a general class of regularizations, and discuss what kind of regularization gives a favorable predictive accuracy. Our main target in this paper is dense type regularizations including \ellp-MKL. According to the recent numeri…
The paper improves model robustness by regularizing posterior differences.
problem Improving model robustness in noisy input scenarios.
method Posterior differential regularization with f-divergence. result Regularizing with f-divergence improves model robustness. Wasserstein distributionally robust optimization (DRO) has recently achieved empirical success for various applications in operations research and machine learning, owing partly to its regularization effect. Although connection between Wasserstein DRO and regularization has been established in several settings, existin…
Study on convergence rates for optimal transport with regularization.
problem Convergence analysis of divergence-regularized optimal transport.
method Novel methodology using quantization and martingale couplings.
result Sharp rates for various divergences and transport costs.
Despite the growing interest in generative adversarial networks (GANs), training GANs remains a challenging problem, both from a theoretical and a practical standpoint. To address this challenge, in this paper, we propose a novel way to exploit the unique geometry of the real data, especially the manifold information. …
New method encodes function preferences into neural nets for better generalization.
problem Challenges in encoding explicit function preferences in neural network training.
method Function-space empirical Bayes (FSEB) regularization.
result FSEB leads to near-perfect semantic shift detection and improved generalization.
Unsupervised representation learning via generative modeling is a staple to many computer vision applications in the absence of labeled data. Variational Autoencoders (VAEs) are powerful generative models that learn representations useful for data generation. However, due to inherent challenges in the training objectiv…
Regularization is typically understood as improving generalization by altering the landscape of local extrema to which the model eventually converges. Deep neural networks (DNNs), however, challenge this view: We show that removing regularization after an initial transient period has little effect on generalization, ev…
Study shows SGD's generalization is not explained by implicit bias.
problem Explaining the generalization ability of overparameterized learning algorithms.
method Revisited Stochastic Convex Optimization with SGD, demonstrating limitations of implicit bias.
result No distribution-independent or distribution-dependent implicit regularizer can explain SGD's generalization.
Exact spectral norm regularization improves neural network generalization.
problem Improving neural network generalization while protecting against noise.
method Exact spectral norm regularization of the Jacobian.
result Improved generalization performance compared to previous methods.
Study reveals the regularization effect of variational distributions in VAEs.
problem Understanding the regularization role of variational distributions in VAEs.
method Analyzed the role of variational family in VAEs and studied the regularization effect on local geometry.
result Uncovered the implicit regularizer in the β-VAE objective and proposed a deterministic autoencoding objective. New insights into how deep models generalize, focusing on matrix factorization.
problem Understanding how deep models generalize and why they work well.
method Using Morse functions and dynamical systems to study implicit regularization.
result Solved a conjecture on implicit regularization in matrix factorization.
New theorem for generalized group sparsity improves consistency and convergence rates.
problem Improving statistical inference in high-dimensional data with element-wise and group-wise sparsity.
method Developed a generalized version of Sparse-Group Lasso and proved a universal theorem for consistency and convergence rates.
result Obtained results on consistency and convergence rates for different forms of double sparsity regularization.
Heavy-tailed regularization improves deep neural network performance.
problem Improving generalization of deep neural networks.
method Introducing Heavy-Tailed Regularization, using differentiable penalty terms and Bayesian statistics.
result Heavy-tailed regularization outperforms conventional regularization techniques.
It is shown that in a tower of coverings the regularized determinant of a generalized Laplacian converges to the L2-determinant. This shows generic nontriviality of analytic torsion or regularized determinants since the L2-counterparts are easier to compute. We further have an "Euler product expansion" for regula…
New regularizer improves neural network robustness and generalization.
problem Ineffective weight decay for networks with homogeneous activation functions.
method Proposes an invariant regularizer to penalize intrinsic weight norms.
result Improves generalization and adversarial robustness on various datasets.
Path regularization improves GFlowNets exploration and generalization.
problem Improving GFlowNets exploration and generalization.
method Path regularization based on optimal transport theory.
result Path regularization enhances GFlowNets to generate more diverse and novel candidates.
New method for training deep neural networks with regularization, converging to better generalization.
problem Improving generalization of deep neural networks through explicit regularization.
method Regularizer Mirror Descent (RMD) method, inspired by convergence properties of stochastic mirror descent (SMD).
result RMD converges to a point close to the minimizer of the cost function, leading to better generalization performance.
Paper finds essential regularity in singular connections.
problem Determining if singularities in connections are removable or essential.
method Introduces RT-equations and a procedure to lift connections to essential regularity.
result A computable procedure to lift connections to essential regularity.
Regularization plays an important role in generalization of deep neural networks, which are often prone to overfitting with their numerous parameters. L1 and L2 regularizers are common regularization tools in machine learning with their simplicity and effectiveness. However, we observe that imposing strong L1 or L2 reg…
We consider the problem of supervised learning with convex loss functions and propose a new form of iterative regularization based on the subgradient method. Unlike other regularization approaches, in iterative regularization no constraint or penalization is considered, and generalization is achieved by (early) stoppin…
Tensor decomposition methods allow us to learn the parameters of latent variable models through decomposition of low-order moments of data. A significant limitation of these algorithms is that there exists no general method to regularize them, and in the past regularization has mostly been performed using bespoke modif…
Dropout improves regularization in flexible models for rare features.
problem Understanding theoretical properties of dropout in generalized linear models.
method Theoretical analysis and application to adaptive smoothing with B-splines.
result Dropout prefers rare features in mean and dispersion parameters.
Proposes RVP to address theoretical concerns of V-REx for OOD generalization.
problem Theoretical concerns about V-REx's motivation and utility.
method Risk Variance Penalization (RVP) modifies V-REx's regularization.
result RVP discovers a robust predictor and finds invariant predictors under certain conditions.
Fiedler regularization uses spectral graph theory to improve neural network performance.
problem Improving neural network performance by penalizing weights based on connectivity.
method Uses the Fiedler value of the neural network's graph as a regularization tool, providing theoretical and computational methods.
result Demonstrates Fiedler regularization's effectiveness in improving neural network performance.
Introduces self-regularization for analyzing learning algorithms.
problem Analyzing and optimizing learning algorithms without explicit regularization.
method Develops a self-regularization framework for learning algorithms.
result Provides statistical analysis and minmax-optimal rates for self-regularized algorithms.
Study of generalized Bishop frames on curves in 4D space.
problem Understanding frames on curves in 4D space.
method Introducing and studying four types of generalized Bishop frames on curves in E4. result Every regular curve in E4 admits all four types of generalized Bishop frames. CASTLE learns causal DAG to improve model generalization.
problem Improving model generalization to out-of-sample data.
method CASTLE learns causal relationships via adjacency matrix embedded in neural network input layers, reconstructing only causal features.
result CASTLE leads to better out-of-sample predictions compared to other regularizers.
Novel regularization for Vision Transformers improves model generalization and sparsity.
problem Improving generalization and sparsity in Vision Transformers.
method Likelihood-guided variational Ising-based regularization.
result Improved generalization and sparsity in Vision Transformers.
The study examines differential smoothness in specific Artin-Schelter regular algebras of dimension 5.
problem Investigating the differential smoothness of Artin-Schelter regular algebras of dimension 5.
method Analyzing the relationship between the number of generators and Gelfand-Kirillov dimension to identify structural obstructions.
result Certain two- and four-generator AS-regular algebras of global dimension five fail to admit a differential calculus, while a five-generator graded Clifford algebra provides a positive example.
New L1 regularization controls neural network generalization error and sparsifies input dimensions.
problem Selecting the optimal number of hidden neurons in neural networks.
method Theoretical analysis of L1 regularization in two-layer neural networks. result Appropriate L1 regularization leads to near minimax optimal generalization risk bounds. The paper studies how different entropic regularizations affect GAN solutions.
problem Improving numerical convergence and sparsity in GAN solutions.
method Entropic regularization of Wasserstein distance and Sinkhorn divergence.
result Entropy regularization promotes sparsity, while Sinkhorn divergence recovers unregularized solution.
We prove regularity results up to the boundary for time independent generalized Maxwell equations on Riemannian manifolds with boundary using the calculus of alternating differential forms. We discuss homogeneous and inhomogeneous boundary data and show 'polynomially weighted' regularity in exterior domains as well.
The paper discusses regularization properties of artificial data for deep learning. Artificial datasets allow to train neural networks in the case of a real data shortage. It is demonstrated that the artificial data generation process, described as injecting noise to high-level features, bears several similarities to e…
We propose and study a general framework for regularized Markov decision processes (MDPs) where the goal is to find an optimal policy that maximizes the expected discounted total reward plus a policy regularization term. The extant entropy-regularized MDPs can be cast into our framework. Moreover, under our framework, …
This work develops a unified framework for RLHF with general f-divergence regularization.
problem Theoretical understanding of general f-divergence regularization in RLHF. method Holistic approach across f-divergence class, two algorithms based on distinct sampling principles. result Provably efficient algorithms with O(logT) regret and O(1/T) sub-optimality gap. The paper shows how policy regularization acts like an adversary to improve robustness.
problem Improving robustness of learned policies in reinforcement learning.
method Using convex duality, the paper characterizes adversarial reward perturbations and provides generalization guarantees.
result Policy regularization acts as an adversary to improve robustness against worst-case reward perturbations.
GPMD solves regularized RL with linear convergence, promoting structural policies.
problem Regularized reinforcement learning to encourage exploration and structural policies.
method Policy mirror descent with generalized convex regularizers and Bregman divergence.
result GPMD converges linearly to the global solution over a wide range of learning rates.
Generic smooth boundaries for isoperimetric regions in 8D manifolds.
problem Understanding boundaries of isoperimetric regions in high-dimensional spaces.
method Generic regularity results for isoperimetric regions in closed Riemannian manifolds of dimension eight.
result Smooth nondegenerate boundaries for isoperimetric regions for generic metrics and volumes.
Many recent successful (deep) reinforcement learning algorithms make use of regularization, generally based on entropy or Kullback-Leibler divergence. We propose a general theory of regularized Markov Decision Processes that generalizes these approaches in two directions: we consider a larger class of regularizers, and…
Extends optimal regularity and compactness to vector bundles over non-Riemannian manifolds.
problem Optimal regularity and compactness for connections on vector bundles.
method Derive RT-equations, establish existence theory, handle curvature up to L1. result Optimal regularity and compactness extended to vector bundles over non-Riemannian manifolds.
Regularization can induce grokking in neural networks, improving generalization.
problem Delayed generalization following overfitting in neural networks.
method Demonstrates that gradient descent with small regularization of model properties induces grokking.
result Regularization can induce grokking, extending previous work on weight decay.
Regularization methods can overregulate, suppressing causal features.
problem Mitigating shortcuts in models exploiting spurious correlations.
method Analysis of regularization methods and their effects on causal features.
result Regularization can overregulate, suppressing causal features.
A fast sketching algorithm solves regularized least squares problems efficiently.
problem Solving large-scale optimization problems with convex or nonconvex regularization.
method Sketching for Regularized Optimization (SRO) algorithm that generates a sketch of the original data matrix and solves the sketched problem.
result General theoretical results for the approximation error between the original and sketched problems, including minimax rates for sparse signal estimation.
Unified framework for analyzing pessimism in off-policy learning with regularized importance sampling.
problem High variance in importance weighting for off-policy learning.
method Unified PAC-Bayesian study of pessimism with regularized importance sampling.
result Derivation of a tractable PAC-Bayesian generalization bound for common importance weight regularizations.
Choquet regularization improves exploration in RL.
problem Improving exploration in reinforcement learning.
method Introducing Choquet regularizers to measure and manage exploration, reformulating RL problems and deriving explicit solutions.
result Explicit optimal distributions and Choquet regularizers for various exploratory samplers.