New algorithms improve robust estimation in contaminated Gaussian models.
problem Simultaneous estimation of location and variance matrix in contaminated Gaussian models.
method Tractable adversarial algorithms with spline discriminators for robust estimation.
result Achieve minimax optimal rates or near-optimal rates under Huber's contamination model.
Robust estimators for Gaussian sparse tasks with optimal error under contamination.
problem Robust mean estimation, PCA, and linear regression in the presence of Huber contamination.
method Novel multidimensional filtering method for sparse regime.
result Optimal error guarantees within constant factors for Gaussian robust k-sparse mean estimation. Study on estimating Gaussian mean with missing data in high dimensions.
problem Estimating Gaussian mean in high dimensions with missing data due to realizable contamination.
method Statistical Query model, Low-Degree Polynomials, PTF tests, and algorithms.
result Established information-computation gap and developed efficient algorithms.
Near-optimal algorithms for mean estimation and linear regression with Gaussian covariates and Huber contamination.
problem Gaussian mean estimation and linear regression with Gaussian covariates in the presence of Huber contamination.
method Near-optimal algorithms with optimal error guarantees, achieving sample complexity n=ildeO(d/ε2) and almost linear runtime. result First sample near-optimal and almost linear-time algorithms with optimal error guarantees for both problems.
Efficiently estimates mean in contaminated Gaussian data with near-optimal sample complexity.
problem Robust mean estimation in the presence of mean-shift contamination.
method First computationally efficient algorithm with near-optimal sample complexity and polynomial-time running.
result Approximates the target mean to any desired accuracy with constant fraction of outliers tolerated.
New winsorized mean improves robustness to up to 50% contamination.
problem Improving robustness of mean estimation in the presence of outliers.
method Outlyingness-induced winsorized mean approach.
result Achieves up to 50% contamination robustness with sub-Gaussian performance.
Robust score matching improves parameter estimation in contaminated data.
problem Parameter estimation in data contaminated by outliers.
method Geometric median of means to develop a robust score matching procedure.
result Consistent parameter estimates in contaminated data settings.
Study tackles contamination and heterogeneity in multi-task learning, improving robustness and personalization.
problem Challenges in integrating related tasks due to contamination and heterogeneity.
method Proposes a filtering-based robust multi-task gradient descent method to estimate global and clean task-specific minimizers.
result Demonstrates improved robustness and personalization compared to existing methods.
Robust covariance testing requires significantly more samples in contaminated data.
problem Testing the covariance matrix of a high-dimensional Gaussian in the presence of contamination.
method We study the problem in the Huber's contamination model, distinguishing between the identity matrix and matrices far from it in Frobenius norm.
result The sample complexity of covariance testing increases dramatically to Ω(d2) in the contaminated setting. Theoretical study on AI models' resilience to data contamination during recursive training.
problem Data contamination in recursive training of generative AI models.
method General framework with minimal assumptions on real data distribution and flexible generative models.
result Contaminated recursive training converges with a rate equal to the minimum of baseline model's rate and contamination fraction.
The paper develops adaptive confidence intervals for Efron's Gaussian two-groups model with unknown contamination.
problem Developing robust uncertainty quantification for Efron's Gaussian two-groups model with unknown contamination fraction.
method The approach involves Fourier-based certification procedures to find minimax-optimal adaptive confidence intervals.
result The minimax-optimal length of adaptive confidence intervals is polynomially worse than when contamination fraction is known.
The paper tackles best arm identification in contaminated bandits with optimal error guarantees and sample complexity.
problem Best arm identification in stochastic bandits with adversarial reward contamination.
method Proposes two algorithms: a gap-based algorithm and a successive elimination-based algorithm for sub-Gaussian bandits.
result Asymptotically optimal sample complexity for both algorithms.
A new robust GP regression algorithm that trims outliers improves model accuracy.
problem Severe bias in GP regression due to data contamination by outliers.
method Iterative trimming of extreme data points.
result Significantly outperforms standard and robust GP variants in most test cases.
Unified framework for robust discriminant analysis overcomes Gaussian assumptions.
problem Challenges in linear and quadratic discriminant analysis with non-Gaussian or contaminated data.
method FEMDA framework considers arbitrary Elliptically Symmetrical (ES) distributions with flexible scale parameters.
result Maximum-likelihood parameter estimation and classification are robust and efficient.
Motivated by applications of bandit algorithms in education, we consider a stochastic multi-armed bandit problem with ε-contaminated rewards. We allow an adversary to give arbitrary unbounded contaminated rewards with full knowledge of the past and future. We impose the constraint that for each time t the…
Improved robust regression for heavy-tailed and contaminated data.
problem Linear regression with heavy-tailed and adversarially contaminated covariates and responses.
method Applying a filtering algorithm to covariates and then using Huber regression, least trimmed squares, or least absolute deviation estimators on the remaining data.
result Near-optimal error rates achieved for the Huber regression estimator.
Paper shows MoM is optimal under adversarial contamination for certain distributions.
problem Optimality of MoM under adversarial contamination.
method Upper and lower bounds for MoM's error under adversarial contamination.
result MoM is (minimax) optimal for distributions with finite variance and infinite variance with finite absolute moments.
This work is motivated by the problem of image mis-registration in remote sensing and we are interested in determining the resulting loss in the accuracy of pattern classification. A statistical formulation is given where we propose to use data contamination to model and understand the phenomenon of image mis-registrat…
New algorithm reduces contamination in supervised learning.
problem Learning with contamination in supervised learning.
method Iterative polynomial filtering.
result Efficient learning of functions with contamination.
Replacing MSE with f-divergence in diffusion models improves robustness under data contamination.
problem Improving robustness of diffusion models under data contamination.
method Replacing MSE with f-divergence in diffusion models.
result Empirical improvement in performance under data contamination.
We propose a novel exponentially-modified Gaussian (EMG) mixture residual model. The EMG mixture is well suited to model residuals that are contaminated by a distribution with positive support. This is in contrast to commonly used robust residual models, like the Huber loss or ℓ1, which assume a symmetric contami…
Neural processes approximate Gaussian process inference, revealing three key costs.
problem Approximating Gaussian process inference with neural processes.
method Bounding KL divergence into three components: label contamination, information bottleneck, and amortization error.
result Characterization of three costs of amortizing Gaussian process inference with neural processes.
A robust loss for anomaly mitigation and unsupervised contamination classification
problem Detecting and mitigating contamination in supervised and unsupervised settings
method Neural Bayesian Anomaly Mitigation (NBAM)
result Recovering the structure of contamination and identifying label-flip pairs
Improved PINNs for solving PDEs with unknown measurement noise.
problem Handling non-Gaussian noise in physics-informed neural networks.
method Jointly train an EBM to learn the correct noise distribution.
result Improved performance in solving PDEs with non-Gaussian noise.
Study develops ensemble machine learning framework for predicting groundwater heavy metal pollution.
problem Statistical complexity and spatial heterogeneity of heavy metal contamination in groundwater.
method Nested cross-validated ensemble machine learning with response transformations (raw, log, Gaussian copula).
result Copula-based models with DBSCAN clustering diagnostics provide the most reliable and interpretable assessments of groundwater contamination.
Study shows efficient algorithms for noiseless linear regression require quadratic sample complexity in contamination rate.
problem Efficient algorithms for noiseless linear regression under Gaussian covariates with oblivious contamination.
method Formal evidence using Statistical Query complexity.
result Any efficient Statistical Query algorithm requires VSTAT complexity at least Ω(d^(1/2)/α^2).
Exponential Lasso improves Lasso's robustness to outliers and heavy-tailed noise.
problem Lasso's sensitivity to outliers and heavy-tailed noise in high-dimensional statistics.
method Integrates an exponential-type loss function into the Lasso framework.
result Achieves strong statistical convergence rates robust to heavy-tailed contamination.
Study minimax robustness in statistical estimation under Wasserstein contamination.
problem Adversarial perturbations in statistical data.
method Developed minimax theory for ℓqr losses under Wasserstein-r contaminations. result Exact minimax risk identified for joint contaminations in location estimation and prediction in linear regression.
SHIFT improves robustness in estimating dose-response functions with heavy-tailed contamination.
problem Outliers bias estimates of average dose-response functions in heavy-tailed data.
method SHIFT combines cross-fit nuisance orthogonalization, Welsch-loss, and defensive OLS refit.
result SHIFT reduces RMSE from 1.03 to 0.33 on localized contamination test.
Generative models can still learn from contaminated data, but with limitations.
problem How much contamination can generative models tolerate?
method Characterized robustness under contaminated enumerations, proving generation is achievable for all countable collections if contamination fraction converges to zero.
result Generation under contamination is achievable for all countable collections if contamination fraction converges to zero, but dense generation is strictly less robust.
New robust discriminant analysis for non-Gaussian data.
problem Classical discriminant analysis struggles with non-Gaussian distributions and contaminated datasets.
method Each data point follows its own ES distribution with arbitrary scale, leading to robust classification.
result Maximum-likelihood estimation and classification are simple, fast, and robust.
Improved robust regression with clean covariates achieves better rates than Huber's model.
problem Robust regression under adaptive contamination of responses with clean covariates.
method Exploiting clean covariates to construct an estimator achieving better rates than Huber's model.
result Improved estimation rate even with constant contamination, achieving consistency.
Proposes R2LDA for improved LDA classifier performance.
problem Poor performance of LDA classifiers in small to comparable data sizes.
method Doubly regularized LDA with automatic parameter selection.
result Consistent and effective performance in noisy test data.
A new robust and flexible classification method for non-Gaussian data.
problem Robustness to scale changes and non-Gaussian distributions in classical discriminant analysis.
method FEMDA uses arbitrary Elliptically Symmetrical distributions and scale parameters for each data point.
result FEMDA is robust to scale changes and outperforms other methods.
New method trims network data to resist adversarial contamination.
problem Adversarial contamination in network data affects statistical and algorithmic performance.
method Proposes a new trimming method operating in model space to address both block and white noise contamination.
result Demonstrates superior performance in simulations compared to direct trimming.
Many machine learning problems can be characterized by mutual contamination models. In these problems, one observes several random samples from different convex combinations of a set of unknown base distributions. It is of interest to decontaminate mutual contamination models, i.e., to recover the base distributions ei…
New algorithm resists contamination in high-dimensional regression with optimal performance.
problem Adversarial and measurement errors in high-dimensional data.
method Adversarial Contamination-resistant Iterative Hard Thresholding (AC-IHT) algorithm.
result Achieves minimax near-optimal estimation and signal-adaptive support recovery.
A low rank matrix X has been contaminated by uniformly distributed noise, missing values, outliers and corrupt entries. Reconstruction of X from the singular values and singular vectors of the contaminated matrix Y is a key problem in machine learning, computer vision and data science. In this paper we show that common…
Theoretical study shows AI models can recover from contaminated training data.
problem Data contamination in AI training can degrade model performance.
method Theoretical analysis and experiments on various data types.
result Models converge to true distribution under mild conditions, with rate dependent on real data fraction.
A contamination in a 3-manifold is an object interpolating between the contact structure and the lamination. Contaminations seem to provide a link between 3-dimensional contact geometry and the classical topology of 3-manifolds, as described in a separate paper. In this paper we deal with contaminations carried by bran…
Bayesian method estimates contamination factor for unsupervised anomaly detection.
problem No good methods for estimating contamination factor in unsupervised anomaly detection.
method Bayesian approach using mixture formulation of anomaly detector outputs.
result Estimated contamination factor distribution is well-calibrated and improves anomaly detection performance.
The paper analyzes how overparameterized models can generalize well in multiclass classification.
problem Generalization in multiclass classification with overparameterized models.
method Survival/contamination analysis framework adapted for multiclass classification.
result Multiclass classification can generalize well even with many classes, unlike regression tasks.
Stochastic gradient descent on l1 loss converges to true parameter in online robust regression.
problem Online robust regression with adversarial noise.
method Stochastic gradient descent on l1 loss.
result Converges to true parameter at a rate independent of contaminated measurements.
New test ensures quality of shared data in machine learning.
problem Ensuring quality of external data in machine learning tasks.
method Distribution-free two-sample testing procedures grounded in conformal outlier detection.
result Identifies valuable external data agents for model personalization.
Study improves generalization bounds for machine learning models in the presence of outliers.
problem Improving model robustness against outliers in machine learning.
method Median-of-Means (MoM) estimator and concentration properties analysis under contamination.
result Derives generalization guarantees for pairwise learning in contaminated data.
Study enhances robustness of In-CVaR based regression models under perturbation and contamination.
problem Enhancing robustness of nonlinear regression models under perturbation and contamination.
method Introduces interval conditional value-at-risk (In-CVaR) and rigorously analyzes its robustness properties under both perturbation and contamination.
result The In-CVaR based estimator is qualitatively robust in terms of the Prokhorov metric if and only if the largest portion of losses is trimmed.
Non-Gaussian component analysis (NGCA) is an unsupervised linear dimension reduction method that extracts low-dimensional non-Gaussian "signals" from high-dimensional data contaminated with Gaussian noise. NGCA can be regarded as a generalization of projection pursuit (PP) and independent component analysis (ICA) to mu…
The paper analyzes how conformal prediction works with contaminated reference data.
problem The impact of contamination on the validity and power of conformal prediction methods.
method The paper analyzes the impact of contamination on the validity of conformal methods and proposes a data-cleaning framework to enhance power.
result The proposed data-cleaning framework can effectively enhance power while maintaining type-I error control.