New framework for robust regularization under uncertain data distributions.
problem Addressing ill-posed inverse problems and statistical estimation under distributional uncertainty.
method Distributionally robust optimal regularization using convex duality.
result Identifies robust regularizers that remain effective under data distributional perturbations.
Study optimal ridge regularization for out-of-distribution prediction.
problem Optimal ridge regularization for predicting out-of-distribution data.
method Established conditions for optimal regularization under covariate and regression shifts, proving monotonic risk in data aspect ratio.
result Negative regularization can be optimal under shifts, even with isotropic or underparameterized training features.
Proposes PER loss to regularize neural network activations to normal distribution.
problem Improving neural network generalization and training speed.
method Regularizes activations to standard normal distribution via projected error function and Wasserstein distance.
result Minimizes Wasserstein distance between activation distribution and standard normal.
Solving logistic regression with L1-regularization in distributed settings is an important problem. This problem arises when training dataset is very large and cannot fit the memory of a single machine. We present d-GLMNET, a new algorithm solving logistic regression with L1-regularization in the distributed settings. …
The paper studies how different entropic regularizations affect GAN solutions.
problem Improving numerical convergence and sparsity in GAN solutions.
method Entropic regularization of Wasserstein distance and Sinkhorn divergence.
result Entropy regularization promotes sparsity, while Sinkhorn divergence recovers unregularized solution.
Method estimates multiple related Gaussian distributions using Laplacian regularization.
problem Jointly estimate multiple related zero-mean Gaussian distributions.
method Laplacian regularized stratified model fitting with hyper-parameters to encourage covariance closeness.
result The method performs well, especially in low data regimes, as demonstrated in finance, radar, and weather.
This work improves distribution recovery from sparse data using Random Forest implicit regularization.
problem Distribution recovery from limited statistics.
method Closed-form estimator for scaled beta distributions, using composite quantile and moment matching.
result Improved classification accuracy through closed-form distribution recovery and implicit regularization.
DAC enhances exploration in reinforcement learning with entropy regularization.
problem Improving exploration efficiency in reinforcement learning.
method Sample-aware entropy regularization using replay buffer action distributions.
result DAC significantly outperforms existing algorithms in reinforcement learning tasks.
Develops a distributed strategy for Pareto optimization of aggregate costs with smoothed regularizers.
problem Optimizing aggregate costs with non-smooth regularizers in a network of agents.
method Distributed strategy using infimal convolution to smooth regularizers, seeking Pareto optimal solution via diffusion.
result Pareto solution of smoothed problem can be made arbitrarily close to original non-smooth problem.
This work improves policy optimization by maximizing entropy of state distribution, leading to better exploration.
problem Lack of exploration in state space when maximizing policy entropy.
method Proposes maximizing the entropy of a lower bound approximation to the state weighting distribution, based on latent space representation.
result Entropy regularization based on marginal state distribution achieves superior state space coverage and better performance in various domains.
We formulate a principle for classification with the knowledge of the marginal distribution over the data points (unlabeled data). The principle is cast in terms of Tikhonov style regularization where the regularization penalty articulates the way in which the marginal density should constrain otherwise unrestricted co…
Deep networks adapt to function regularity and data distribution.
problem Understanding deep learning's adaptability to function regularity and data distribution.
method Developed nonparametric approximation and estimation theories for a broad class of functions using deep ReLU networks.
result Deep neural networks are adaptive to different regularity of functions and nonuniform data distributions.
Paper develops a new method for distribution regression with indefinite kernels.
problem Distribution regression with indefinite kernels.
method Coefficient-based regularized distribution regression with two-stage sampling.
result Optimal learning rates derived for the algorithm under mild conditions.
Study of entropy-regularized LQG MFGs with exploratory actions.
problem Optimizing multi-population mean field games with entropy regularization.
method Introduced exploratory actions and derived optimal action distributions.
result Optimal action distributions lead to ε-Nash equilibria in finite-population MFGs.
The paper explores optimal regularizers for data sources, linking them to star bodies.
problem Understanding optimal regularizers for data sources.
method Investigates optimal regularizers for data distributions using star bodies and dual Brunn-Minkowski theory.
result Identifies optimal regularizers and assesses amenability to convex regularization.
Study on GANs learning distributions, deriving rates and regularization.
problem Learning distributions with GANs.
method Analysis of GANs through regularization theory.
result Optimal rates for distribution estimation under adversarial framework.
Choquet regularization improves exploration in RL.
problem Improving exploration in reinforcement learning.
method Introducing Choquet regularizers to measure and manage exploration, reformulating RL problems and deriving explicit solutions.
result Explicit optimal distributions and Choquet regularizers for various exploratory samplers.
This work shows how exploiting gradient alignment can improve distributed and federated learning performance.
problem Misalignment of gradients across clients in distributed and federated learning.
method Utilizing implicit regularization through a novel GradAlign algorithm that induces gradient alignment with large mini-batches.
result Improvements in test accuracies and generalization performance.
We align distributional data using regularized Wasserstein means.
problem Aligning distributional data from different domains.
method Regularized Wasserstein means with variational transportation.
result Sparse representation captures desired properties and reduces mapping cost.
Improves autoencoder reconstruction quality by approximating latent space with adversarial learning.
problem Ambiguity in autoencoder reconstructions and difficulty in matching true distribution.
method Adversarially Approximated Autoencoder (AAAE) using GAN for flexible latent space approximation.
result Generates more faithful reconstructions and maintains latent manifold structure.
Paper proposes a new regularization method to prevent model degradation under distribution shifts.
problem Model performance degrades under distribution shifts.
method Supervised contrastive learning with heterogeneous similarity.
result The proposed method outperforms existing regularization methods on benchmark datasets.
No regularization needed for InLDL, achieving efficient and effective model.
problem InLDL struggles with performance degradation due to missing degrees.
method Proposes a model that uses label distribution as a prior, implicitly regularizing the learning process.
result Achieves competitive performance without explicit regularization.
Proposes a method to ensure latent space of AutoEncoders is Gaussian without randomness.
problem Ensuring latent space of AutoEncoders is Gaussian without randomness.
method Directly optimize agreement between empirical distribution function and desired CDF for chosen properties.
result Ensures latent space of AutoEncoders is Gaussian without randomness.
For a 2-dimensional non-flat spray we associate a Berwald frame and a 3-dimensional distribution that we call the Berwald distribution. The Frobenius integrability of the Berwald distribution characterises the Finsler metrizability of the given spray. In the integrable case, the sought after Finsler function is pro…
Monotonic relationship found between in-distribution and out-of-distribution performance.
problem Understanding performance of machine learning models under distribution shifts.
method Analyzing ridge-regularized models and linear inverse problems under covariate shift.
result Monotonic relationship between in-distribution and out-of-distribution performance for certain models.
Unified framework for constructing nonconvex sparse recovery methods.
problem Constructing valid nonconvex regularization functions remains open.
method Unified framework based on probability density function, using Weibull distribution.
result New nonconvex sparse recovery method based on Weibull distribution.
Positive mass theorem for asymptotically flat manifolds with non-negative distributional scalar curvature
problem Positive mass theorem
method Ricci flow smoothing
result Asymptotically flat manifolds with non-negative ADM mass
The paper introduces a novel method for training neural network Stein critics with staged L2-regularization.
problem Learning to differentiate model distributions from observed data in high-dimensional settings.
method Developed a novel staging procedure for L2 regularization over training time, leveraging the advantages of highly-regularized training at early times. result Theoretical guarantees and empirical validation show that the method improves the approximation of the training dynamic by the kernel optimization, leading to faster convergence and better performance.
Gaussian Processes (GPs) are a popular approach to predict the output of a parameterized experiment. They have many applications in the field of Computer Experiments, in particular to perform sensitivity analysis, adaptive design of experiments and global optimization. Nearly all of the applications of GPs require the …
We propose a practical method for L0 norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known …
Paper proposes SinkhornDRL for distributional RL using Sinkhorn divergence and regularized Wasserstein loss.
problem Improving distributional reinforcement learning by minimizing Bellman return distribution differences.
method Introduces SinkhornDRL, a distributional RL algorithm using Sinkhorn divergence and regularized Wasserstein loss.
result SinkhornDRL consistently outperforms or matches existing algorithms on Atari games, especially in multi-dimensional reward settings.
New sampling method using regularized Wasserstein proximal for Gibbs distributions.
problem Sampling from Gibbs distributions with numerical stability and efficiency.
method Preconditioned regularized Wasserstein proximal operator.
result Discrete-time convergence analysis and explicit bias characterization.
The paper connects three machine learning methods to reduce generalization errors.
problem Reducing generalization errors in machine learning models.
method Distributionally robust optimization, Bayesian methods, and regularization.
result Machine learning models can be characterized using distributional uncertainty and robustness measures.
Data-driven Distributionally Robust Optimization (DD-DRO) via optimal transport has been shown to encompass a wide range of popular machine learning algorithms. The distributional uncertainty size is often shown to correspond to the regularization parameter. The type of regularization (e.g. the norm used to regularize)…
New equivalences found between subsampling and ridge regularization methods.
problem Establishing precise structural and risk equivalences between subsampling and ridge regularization.
method Proved structural and risk equivalences between subsample ridge estimators and different ridge regularization levels and subsample aspect ratios.
result Optimally tuned ridge regression exhibits a monotonic prediction risk in the data aspect ratio.
Mitigates gender bias amplification in model predictions.
problem Gender bias amplification in model predictions.
method Posterior regularization to mitigate bias.
result Almost removes gender bias amplification in model predictions.
dpVAEs improve VAEs by decoupling representation and generation.
problem VAEs struggle with both representation learning and sample generation.
method Introduce decoupled priors (dpVAEs) that separate representation and generation spaces.
result dpVAEs enable regularization without compromising sample generation.
Probabilistic graphical models compactly represent joint distributions by decomposing them into factors over subsets of random variables. In Bayesian networks, the factors are conditional probability distributions. For many problems, common information exists among those factors. Adding similarity restrictions can be v…
New bounds for low-regularity Riemannian metrics defined via distributional curvature.
problem Establishing curvature bounds for Riemannian metrics of low regularity.
method Introducing a distributional version of sectional curvature for C1 and C0 metrics. result New bounds for low-regularity metrics recover classical bounds in Alexandrov spaces.
New regularizer for machine learning using private data.
problem Machine learning with private data.
method Distributionally-robust optimization with locally-differentially-private datasets.
result New regularizer for training linear regression models.
The paper analyzes the complexity of manifold regularization methods.
problem Understanding the complexity of manifold regularization in semi-supervised learning.
method The paper derives sample complexity bounds and Rademacher bounds for semi-supervised methods.
result The semi-supervised method can only have a constant improvement, ignoring logarithmic terms.
DRIVE improves IV estimation by accounting for distributional uncertainties.
problem Challenges in IV estimation due to untestable model assumptions and poor finite sample properties.
method DRIVE is a distributionally robust IV estimation method that minimizes a square root TSLS objective with a Wasserstein ambiguity set.
result DRIVE achieves consistency without requiring regularization parameter to vanish, ensuring robustness to distributional uncertainties.
Convex surfaces derived from specific Riemannian manifolds with high regularity.
problem Proving convexity of surfaces derived from Riemannian manifolds.
method Analyzing solutions to the very weak Monge-Ampère equation.
result Proved convexity of weakly regular surfaces with nonnegative intrinsic curvature.
Unified techniques improve stability and replicability in changing data.
problem Concept drift in data generating distribution.
method Removing hidden confounding and causal regularization.
result Improves stability, replicability, and robustness in heterogeneous data.
Method introduces topological regularization using information filtering networks.
problem Sparse probabilistic modeling and multicollinear regression.
method Topological regularization via information filtering network.
result Direct application to L0-norm regularized problems. New technique debiases distributed optimization, improving convergence rate.
problem Bias in local estimates limits effectiveness of distributed second order optimization.
method Surrogate sketching and scaled regularization to eliminate bias.
result The debiased local estimates lead to faster convergence in distributed optimization.
New autoencoder framework learns structured latent priors.
problem Learning autoencoders with flexible priors.
method Relational regularization on latent prior, scalable algorithms.
result RAE outperforms existing autoencoders in image generation.
This paper studies nonholonomic constraints in Hamiltonian systems, deriving equations and theorems.
problem Analyzing nonholonomic constraints in Hamiltonian systems.
method Deriving distributional RCH systems, geometric constraint conditions, and Hamilton-Jacobi theorems.
result Derives precise geometric constraint conditions and Hamilton-Jacobi theorems for nonholonomic systems.