Label smoothing improves model performance even with noisy labels.
problem Mitigating label noise in deep learning models.
method Examined label smoothing as a technique to cope with label noise and compared it to loss-correction methods.
result Label smoothing is competitive with loss-correction techniques under label noise and beneficial for distillation from noisy data.
Study revisits AdaGrad convergence with relaxed noise assumptions.
problem Non-convex smooth optimization problems with general noise.
method General noise model with function value gap and gradient magnitude control.
result Probabilistic convergence rate of ( ilde{\mathcal{O}}(1/\sqrt{T})) under general noise.
New method enhances neural network robustness against adversarial attacks.
problem Enhancing neural network robustness against adversarial attacks.
method Variational framework with per-sample noise level selector.
result Enhanced empirical robustness and certified robustness.
Smooths GPS data with splines for noisy, irregularly sampled data.
problem Noisy, irregularly sampled GPS data with non-Gaussian noise.
method Smoothing splines with chosen spline order and tension parameter, allowing for non-Gaussian noise and outliers.
result Effective smoothing and interpolation of GPS data.
Paper analyzes convergence of stochastic methods under heavy-tailed noise.
problem Analyzing convergence of stochastic methods under heavy-tailed noise.
method Investigates vanilla and clipped stochastic subgradient descent methods.
result Demonstrates convergence properties under sub-Weibull and p-BCM noise assumptions.
GNC smooths loss function for large-batch SGD, improving generalization.
problem Extremely large-batch SGD leads to poor generalization and converges to sharp minima.
method Gradient noise convolution (GNC) smooths loss function by convolving gradient noise with the loss function.
result GNC achieves state-of-the-art generalization performance for large-scale deep neural networks.
New method handles correlated and repeated measurements using smoothed multivariate square-root Lasso.
problem Handling correlated and repeated measurements with complex noise structure.
method Proposes a concomitant estimator that uses non-averaged measurements and leverages smoothing theory for optimization.
result Demonstrates practical benefits on various datasets (toy, simulated, real neuroimaging).
Randomized smoothing reduces accuracy in ML models, especially at higher noise levels.
problem Adversarial attacks on ML models, especially randomized smoothing's accuracy drop.
method Theoretical and empirical analysis of randomized smoothing's effect on feasible hypotheses space.
result For some noise levels, randomized smoothing shrinks the set of feasible hypotheses, leading to accuracy drops.
Hierarchical randomized smoothing improves model robustness for complex data.
problem Certifying robustness on complex data (e.g. images, graphs) is challenging.
method Add random noise to a randomly selected subset of entities in a hierarchical manner.
result Hierarchical randomized smoothing yields stronger robustness guarantees with high accuracy.
Certified robustness for ImageNet models with randomized smoothing.
problem Creating robust models against adversarial attacks.
method Randomized smoothing with Gaussian noise.
result Certified top-1 accuracy of 49% on ImageNet under small ℓ 2 \ell_2 ℓ 2 perturbations. We consider the non-parametric regression problem under Huber's ε ε ε -contamination model, in which an ε ε ε fraction of observations are subject to arbitrary adversarial noise. We first show that a simple local binning median step can effectively remove the adversary noise and this median estimator is minimax optimal up t…
New adaptive algorithm optimizes functions with minimal smoothness and noise.
problem Optimizing functions with unknown noise levels and minimal local smoothness.
method Simple, parameter-free approach that adapts to unknown noise levels and local smoothness.
result First algorithm that naturally adapts to unknown noise levels and local smoothness.
Regularization improves robustness of smoothed classifiers.
problem Certifying robustness of smoothed classifiers.
method Regularizing prediction consistency over Gaussian noise.
result Significantly improved certified robustness with less training costs.
New method for faster convergence in non-convex optimization with unbounded smoothness.
problem Finding first-order stationary points of non-convex functions with unbounded smoothness.
method Developed a stopped analysis technique to prove convergence rates for ( L 0 , L 1 ) (L_0,L_1) ( L 0 , L 1 ) -smooth functions. result Achieved O ( p o l y log ( T ) T ) \mathcal{O}(\frac{\mathrm{poly}\log(T)}{\sqrt{T}}) O ( T poly l o g ( T ) ) convergence rates without uniform noise bounds. Paper studies Adam's convergence under relaxed assumptions, proving a rate of O(poly(log T)/sqrt(T)).
problem Understanding Adam's convergence in non-convex, stochastic optimization with unbounded gradients and noise.
method Introduced a comprehensive noise model and used it to prove Adam's convergence rate.
result Adam finds a stationary point with a rate of O(poly(log T)/sqrt(T)) in high probability.
VAE with noise model learns smoothed densities without seeing noisy data.
problem Learning smoothed densities with noisy data.
method Imaginary noise model in variational autoencoders (σ-VAE).
result All σ-VAEs are equivalent via β-VAE expansion.
New framework assesses regularization norms in ill-posed problems, revealing L2 instability and proposing adaptive fractional RKHS solutions.
problem Comparative analysis of regularization norms in ill-posed problems.
method Small noise analysis framework for Tikhonov and RKHS regularizations.
result Optimal convergence rates achieved with adaptive fractional RKHS, but hyper-parameters decay too fast.
Unified framework for distributed compressed SGD under ( L 0 , L 1 ) (L_0, L_1) ( L 0 , L 1 ) -smoothness.
problem Understanding the joint effect of batch noise, adaptivity, and compression in distributed stochastic optimization.
method Developed a unified theoretical framework using SDEs that incorporate curvature-dependent terms.
result Normalizing updates in DCSGD stabilizes convergence, with normalization degree determined by noise structure and landscape regularity.
Random smoothing struggles to certify high-dimensional image robustness.
problem Certifying adversarial robustness for high-dimensional images with p > 2 p>2 p > 2 . method Analysis of random smoothing for ℓ p \ell_p ℓ p robustness, focusing on ℓ ∞ \ell_\infty ℓ ∞ . result Noise distribution required for ℓ p \ell_p ℓ p robustness must have high variance, leading to trivial classifiers. This work addresses various open questions in the theory of active learning for nonparametric classification. Our contributions are both statistical and algorithmic: -We establish new minimax-rates for active learning under common \textit{noise conditions}. These rates display interesting transitions -- due to the inte…
New algorithm speeds up RNN time series prediction by filtering noise.
problem Predicting smooth trajectories from noisy time series data.
method Analyzed RNN dynamics to propose an efficient noise filtering algorithm.
result Significant speedup in predictive process without accuracy loss.
Optimized method tackles convex optimization with heavy-tailed noise.
problem Convex optimization problems with noisy gradients.
method Vanilla stochastic proximal subgradient method without gradient clipping or normalization.
result Achieves optimal complexity for various convex optimization types under heavy-tailed noise.
Kernel regression predicts graph signals in noisy environments.
problem Predicting smooth graph signals in the presence of sparse noise.
method Kernel regression with ℓ 1 \ell_1 ℓ 1 -norm and ℓ 2 \ell_2 ℓ 2 -norm optimization using IRLS. result Efficacy demonstrated on real-world temperature data.
STAG injects noise into graph neural networks to improve performance.
problem Graph neural networks suffer from over-smoothing and limited discrimination.
method Introduces a stochastic aggregation framework (STAG) with adaptive noise injection.
result STAG models correct both over-smoothing and discrimination issues.
Solves learning halfspaces with Massart noise for log-concave distributions.
problem Learning halfspaces with Massart noise in distribution-specific PAC model.
method Identifies a smooth non-convex surrogate loss and uses SGD to solve the learning problem.
result First computationally efficient algorithm for learning halfspaces with Massart noise for a broad family of distributions.
Graph attention is not always beneficial; conditions for perfect node classification are identified.
problem Understanding when graph attention mechanisms improve node classification performance.
method Theoretical analysis using Contextual Stochastic Block Models (CSBMs).
result Graph attention mechanisms are more effective when structure noise exceeds feature noise, and simpler graph convolution operations are better when feature noise predominates.
DP-LSSGD improves privacy-preserving ML models by smoothing out noise.
problem Privacy-preserving ML models have lower utility than non-private ones.
method DP-LSSGD uses Laplacian smoothing to improve utility of DP-SGD.
result DP-LSSGD achieves the same DP guarantee as DP-SGD but with better stability and generalization.
New methods solve optimization problems with heavy-tailed noise, improving upon existing complexity bounds.
problem Optimization problems with heavy-tailed noise and weakly average smoothness.
method Normalized stochastic first-order methods with Polyak, multi-extrapolated, and recursive momentum.
result First-order oracle complexity results for finding approximate stochastic stationary points under heavy-tailed noise.
Label smoothing improves generalization by controlling generalization loss.
problem Lack of mathematical understanding of label smoothing's effectiveness.
method Proposed a theoretical framework to show how label smoothing controls generalization loss in the label noise setting.
result Predicted an optimal label smoothing point that minimizes generalization loss.
Polarimetric Synthetic Aperture Radar (PolSAR) images are establishing as an important source of information in remote sensing applications. The most complete format this type of imaging produces consists of complex-valued Hermitian matrices in every image coordinate and, as such, their visualization is challenging. Th…
This work extends score-based methods to binary data on the Boolean hypercube.
problem Learning and sampling binary data on the Boolean hypercube.
method Adopting Bernoulli noise as a smoothing device, deriving a TMF-like expression for the optimal denoiser, and using a Langevin-like sampler.
result The method successfully samples noisy binary data and reduces effective noise through multiple measurements.
A new algorithm improves both computational efficiency and statistical optimality for robust low-rank matrix and tensor estimation.
problem Challenges in low-rank matrix estimation under heavy-tailed noise, both computationally and statistically.
method Riemannian sub-gradient (RsGrad) algorithm, which is computationally efficient and statistically optimal.
result RsGrad achieves linear convergence and statistical optimality for robust loss functions under Gaussian and heavy-tailed noise.
Prediction of dynamical time series with additive noise using support vector machines or kernel based regression has been proved to be consistent for certain classes of discrete dynamical systems. Consistency implies that these methods are effective at computing the expected value of a point at a future time given the …
Method reweights instances and classes to improve robustness in noisy data.
problem Improving deep learning performance in the presence of label noise.
method Formulates constrained optimization problems to assign importance weights to instances and class labels.
result Significant performance gains observed in benchmark datasets with label noise.
New framework improves adversarial robustness certification for various perturbations.
problem Certifying robustness against adversarial attacks in deep learning models.
method Unified functional optimization approach with non-Gaussian smoothing noise for multiple types of attacks.
result Achieves better certification results and identifies key trade-offs between accuracy and robustness.
Adapts SGD to noise and problem specifics for faster convergence.
problem Minimizing smooth, strongly-convex functions with varying noise and problem constants.
method Adaptive SGD with exponentially decreasing step-sizes, Nesterov acceleration, and stochastic line-search.
result Achieves near-optimal convergence rates without knowing noise or problem specifics.
Nonparametric density deconvolution and denoising using simulation-based inference
problem Learning latent signals and their distributions in the presence of measurement noise
method Convolutional maximum mean discrepancy (convMMD) loss and likelihood-free framework
result Learn a latent generative model matching observed data distribution
Paper explores whether gradient normalization can replace clipping for SGD in heavy-tailed noise.
problem Ensuring convergence of SGD in heavy-tailed noise.
method Revisits gradient clipping and normalization, proving their sufficiency and effectiveness.
result Gradient normalization alone is sufficient for nonconvex SGD convergence under smoothness assumptions.
In Deep Learning, Stochastic Gradient Descent (SGD) is usually selected as a training method because of its efficiency; however, recently, a problem in SGD gains research interest: sharp minima in Deep Neural Networks (DNNs) have poor generalization; especially, large-batch SGD tends to converge to sharp minima. It bec…
Improved k-NN active learning with local smoothness assumption.
problem Active learning convergence rates under smoothness assumptions.
method Designing an active learning algorithm with better convergence rate using local smoothness assumption for k-NN.
result Better convergence rate than in passive learning.
RESTA defends LLMs against jailbreaking attacks by adding random noise to embeddings.
problem Vulnerability of LLMs to jailbreaking attacks that generate harmful outputs.
method Adds random noise to embedding vectors and aggregates during token generation.
result RESTA achieves superior robustness versus utility tradeoffs compared to baseline defenses.
Hidden cost: Smoothing shrinks decision boundaries, affecting class-wise accuracy.
problem The fragility of machine learning models and the need for robustness verification.
method Randomized smoothing approach to achieve statistical robustness.
result Smoothed classifiers' decision boundaries shrink, leading to class-wise accuracy disparity.
Novel active learning algorithm with improved convergence rate under local smoothness condition.
problem Improving convergence rates in active learning under specific smoothness assumptions.
method Developed a novel active learning algorithm with a rate of convergence better than in passive learning, using a local smoothness assumption for k-nearest neighbors.
result The algorithm achieves a better convergence rate than passive learning algorithms, avoiding strong density assumptions.
Paper introduces a new regularization method for kernel gradient descent learning.
problem Preventing overfitting in kernel gradient descent learning.
method Random smoothing regularization as novel convolution-based smoothing kernels.
result Optimal convergence rates achieved in various function spaces.
Generalizes smoothness conditions for optimization methods.
problem Optimization under non-uniform smoothness conditions.
method Develops a new analysis technique for bounding gradients.
result Obtains convergence rates for gradient descent and Nesterov's method.
New method improves robustness of large models without sacrificing accuracy.
problem Improving robustness of large pre-trained models without accuracy loss.
method Multi-scale diffusion denoised smoothing, selectively applying smoothing at multiple noise scales.
result Strong certified robustness at high noise levels with accuracy close to non-smoothed classifiers.
A new robust gradient descent method improves generalization efficiency.
problem Improving off-sample generalization of learning algorithms under heavy-tailed data.
method Smoothed multiplicative noise applied to observations before constructing a sum of soft-truncated gradient coordinates.
result The proposed method achieves competitive theoretical guarantees and efficient generalization over a wide class of data distributions.
New algorithms SVCA and SSPA improve robustness to noise in nonnegative matrix factorization.
problem Estimating vertices from noisy data points in convex hull.
method Smoothed VCA (SVCA) and Smoothed SPA (SSPA) algorithms.
result Improved robustness to noise compared to existing methods.