This paper investigates differentially private analysis of distance-based outliers. The problem of outlier detection is to find a small number of instances that are apparently distant from the remaining instances. On the other hand, the objective of differential privacy is to conceal presence (or absence) of any partic…
Outliers are ubiquitous in modern data sets. Distance-based techniques are a popular non-parametric approach to outlier detection as they require no prior assumptions on the data generating distribution and are simple to implement. Scaling these techniques to massive data sets without sacrificing accuracy is a challeng…
Enhances clustering for functional data, robust to outliers.
problem Challenges of clustering infinite-dimensional functional data and outlier sensitivity.
method Extends OCLUST algorithm to handle functional data, trimming outliers.
result Strong performance in clustering and outlier identification on simulated and real-world datasets.
New framework detects outliers in non-IID categorical data.
problem Existing outlier detection methods fail in non-IID data.
method Value-value graph-based representation and outlierness propagation.
result Significant improvement in AUC on complex data sets.
Chain-ladder reserving is sensitive to outliers, leading to unreliable estimates.
problem Sensitivity of loss reserving techniques to outliers.
method Derivation of impact functions for reserves and mean squared errors of prediction under Mack's Model.
result Impact of outliers varies widely in a loss triangle and depends on other cells.
Improved LDA with capped l_{2,1}-norm reduces outlier sensitivity.
problem Outliers and noise sensitivity in classical LDA.
method Introducing capped l_{2,1}-norm and proposing CLDA.
result CLDA effectively removes outliers and suppresses noise.
New method detects business-relevant outliers in e-commerce conversion rates.
problem Identifying outliers in e-commerce conversion rate data.
method A novel unsupervised fluid IQR method that adjusts sensitivity based on platform activity.
result Fluid IQR method outperforms existing methods in business-relevance and robustness.
Paper proposes methods to make OT robust to outliers.
problem Optimal transport is sensitive to outliers.
method Detect outliers using adversarial training, adjust transport cost based on classifier predictions.
result Outliers are detected and do not affect transport in experiments.
New method SP-PPCA reduces outlier impact in PCA.
problem Outliers make standard PCA and PPCA less robust.
method Integrates self-paced learning into PPCA, using iterative optimization.
result SP-PPCA effectively reduces or eliminates outlier impact.
Proposes AE for robust PCA, improving robustness to outliers.
problem PCA's sensitivity to outliers.
method Angular Embedding (AE) and Truncated Angular Embedding (TAE).
result AE/TAE outperforms state-of-the-art RPCA methods.
Notwithstanding the popularity of conventional clustering algorithms such as K-means and probabilistic clustering, their clustering results are sensitive to the presence of outliers in the data. Even a few outliers can compromise the ability of these algorithms to identify meaningful hidden structures rendering their o…
Proposes robust ABC method for outlier detection.
problem Outliers sensitivity in ABC methods.
method γ-divergence estimator with redescending property.
result Significantly higher robustness than existing methods.
Plain vanilla K-means clustering has proven to be successful in practice, yet it suffers from outlier sensitivity and may produce highly unbalanced clusters. To mitigate both shortcomings, we formulate a joint outlier detection and clustering problem, which assigns a prescribed number of datapoints to an auxiliary outl…
A new method detects outliers using ensembles of Dirichlet process mixtures.
problem Challenges in unsupervised outlier detection using Dirichlet process mixtures.
method Ensembles of Dirichlet process Gaussian mixtures with random subspace and subsampling.
result Empirically outperforms existing approaches in unsupervised outlier detection.
New robust bandit algorithm for clinical trials reduces sensitivity to outlier data.
problem Adaptive clinical trials need a robust bandit algorithm to handle outlier data.
method Proposes a new robustness criterion and modifies BESA algorithm for bandit problems.
result Empirical evaluation shows improved performance compared to standard bandit algorithms.
New robust MPCA method handles casewise and cellwise outliers in tensor data.
problem Outliers, especially casewise and cellwise, affect the performance of standard MPCA.
method Uses a single loss function to reduce the influence of both types of outliers and missing values.
result The new method improves robustness and performance in tensor data analysis.
Robust TOT regression method handles outliers in tensor data.
problem Outliers in tensor data affect standard TOT regression.
method ROTOT method using a single loss function for outliers and robust MPCA for predictor.
result ROTOT method reduces influence of both casewise and cellwise outliers.
Paper introduces LR to generate synthetic data with privacy protection.
problem Privacy concerns limit the use of sensitive datasets.
method Local Resampler (LR) using k-nearest neighbors algorithm.
result LR effectively mitigates outlier-driven disclosure risks.
Improves data normality with robust transformations.
problem Skewed data distribution.
method Modified Box-Cox and Yeo-Johnson transformations with robust parameter estimation.
result Transformed data approximates normality in the center with outliers.
Traditional dictionary learning methods are based on quadratic convex loss function and thus are sensitive to outliers. In this paper, we propose a generic framework for robust dictionary learning based on concave losses. We provide results on composition of concave functions, notably regarding super-gradient computati…
Robust clustering algorithm for datasets with outliers.
problem Clustering with arbitrary outliers.
method Spectral clustering with a rounding scheme on a Gaussian kernel matrix.
result Misclassification error decays exponentially with signal-to-noise ratio.
The support vector machine (SVM) is one of the most successful learning methods for solving classification problems. Despite its popularity, SVM has a serious drawback, that is sensitivity to outliers in training samples. The penalty on misclassification is defined by a convex loss called the hinge loss, and the unboun…
In this paper, a scale mixture of Normal distributions model is developed for classification and clustering of data having outliers and missing values. The classification method, based on a mixture model, focuses on the introduction of latent variables that gives us the possibility to handle sensitivity of model to out…
Many traditional methods for identifying changepoints can struggle in the presence of outliers, or when the noise is heavy-tailed. Often they will infer additional changepoints in order to fit the outliers. To overcome this problem, data often needs to be pre-processed to remove outliers, though this is difficult for a…
Proposes MSD-Kmeans for efficient outlier detection.
problem Detecting unusual records in noisy data.
method Combines MSD and K-means for more accurate outlier detection.
result MSD-Kmeans achieves highest precision, accuracy, and F-measure.
Ordinary least square (OLS) estimation of a linear regression model is well-known to be highly sensitive to outliers. It is common practice to (1) identify and remove outliers by looking at the data and (2) to fit OLS and form confidence intervals and p-values on the remaining data as if this were the original data col…
Truncated CauchyNMF robustly learns subspaces from noisy data.
problem Outliers in non-negative matrix factorization (NMF) cause failure.
method Proposes Truncated CauchyNMF loss to handle outliers.
result Theoretical analysis and experimental validation show Truncated CauchyNMF's robustness.
Develops robust persistence diagrams using kernel methods.
problem Persistence diagrams are sensitive to data perturbations.
method Constructs robust persistence diagrams from superlevel filtrations of robust density estimators using reproducing kernels.
result Robust persistence diagrams are consistent estimators in bottleneck distance.
Robust PCA reduces to power iterations for outlier-resilient feature extraction.
problem Sensitivity of PCA to non-Gaussian samples and outliers.
method Robust formulation of PCA based on maximum correntropy criterion.
result MCPI reduces to power iterations, making PCA more robust to outliers.
A new robust Wasserstein distance is proposed to handle outliers in probability distributions.
problem Outliers in probability distributions make Wasserstein distances sensitive and impractical.
method Introduces a new outlier-robust Wasserstein distance Wpε. result Achieves strong robust estimation guarantees under the Huber ε-contamination model. A new framework detects adverse dataset shifts using outlier scores.
problem False alarms in dataset shift tests.
method Outlier scores to compare contamination rates at varying thresholds.
result Reduces the sensitivity to minor differences in predictive performance.
Mean embeddings provide an extremely flexible and powerful tool in machine learning and statistics to represent probability distributions and define a semi-metric (MMD, maximum mean discrepancy; also called N-distance or energy distance), with numerous successful applications. The representation is constructed as the e…
Robust GW distance improves graph data alignment.
problem Outliers in GW distance lead to inaccurate comparisons.
method Optimistically perturbed marginal constraints within a Kullback-Leibler divergence-based ambiguity set.
result RGW reduces inaccuracies in graph data alignment.
New method robustifies topological data analysis against outliers.
problem Outliers make topological data analysis unstable.
method Proposed a robust distance function (MoM Dist) for persistent homology.
result MoM Dist sublevel filtrations and weighted filtrations are consistent estimators in adversarial settings.
Novel method measures DNN sensitivity to perturbations.
problem Vulnerability of DNNs to adversarial examples.
method Perturbation manifold and influence measure.
result Demonstrated usefulness for model building tasks.
DORO improves DRO's performance and stability in tasks with subpopulation shift.
problem DRO's poor performance and instability in tasks with subpopulation shift.
method DORO, a refined risk function that prevents overfitting to outliers.
result DORO improves DRO's performance and stability on large modern datasets.
A new robust metric compares distributions more accurately than existing methods.
problem Sensitivity to outliers and sampling discrepancy in Wasserstein distances.
method Introducing k-RPW, a partial p-Wasserstein distance.
result k-RPW converges faster to true distance and is more robust to outliers.
A new robust scaling approach improves downstream metabolomics analysis.
problem Challenges in choosing scaling techniques for metabolomics data.
method Introduces a weighted scaling approach robust to outliers.
result The proposed method outperforms traditional scaling techniques in both outlier-free and outlier-present datasets.
Bayesian optimization has recently attracted the attention of the automatic machine learning community for its excellent results in hyperparameter tuning. BO is characterized by the sample efficiency with which it can optimize expensive black-box functions. The efficiency is achieved in a similar fashion to the learnin…
A new PCA method using Tℓ1-norm outperforms existing methods.
problem Outliers and noise sensitivity in classical PCA.
method PCA based on Tℓ1-norm maximization. result The method outperforms PCA-ℓp, ℓpSPCA, and PCA in numerical experiments. Unbalanced COOT improves feature alignment robustly to outliers.
problem Optimal transport methods are sensitive to outliers in real-world data.
method COOT infers alignment between features and samples, unbalanced COOT adds robustness.
result Unbalanced COOT is robust to noise in real-world datasets.
Unified framework TERM improves fairness and robustness.
problem Outliers and subgroup fairness in empirical risk minimization.
method Unified framework TERM with a hyperparameter tilt.
result TERM improves fairness and robustness.
Proposes a robust portfolio method for large asset universes.
problem Outliers in return data affect traditional portfolio optimizations.
method Robust PCA, shrinkage estimation, and adaptive portfolio weights.
result Superior portfolio performance in numerical and empirical tests.
Proposes MPCA for robust PCA using mode estimation.
problem Outliers sensitivity in PCA.
method Modal Principal Component Analysis (MPCA) based on mode estimation.
result MPCA shows advantages over conventional methods.
Stabilizes online learning by using weighted reservoir sampling.
problem Real-world deployment sensitivity to outliers causes low accuracy in final solutions.
method Weighted reservoir sampling to stabilize ensemble model without additional data passes.
result Risk of ensemble classifier is bounded with respect to the underlying online learning method's regret.
Online TERM improves robustness and fairness in streaming data.
problem Streaming data's lack of worst-case fairness and robustness in ERM.
method Proposes an online TERM formulation to balance average-case accuracy with worst-case fairness and robustness.
result Negative tilting effectively suppresses outlier influence, positive tilting improves recall with minimal precision loss.
Hölder-Bayes robustly infers model parameters and contamination levels.
problem Robustness to data contamination in Bayesian inference.
method Introduces Hölder-Bayes framework for joint inference of model parameters and contamination proportion using Hölder divergence.
result Hölder-Bayes framework provides robust parameter inference, contamination-level recovery, and uncertainty-aware outlier detection.
Bayesian learning is often hampered by large computational expense. As a powerful generalization of popular belief propagation, expectation propagation (EP) efficiently approximates the exact Bayesian computation. Nevertheless, EP can be sensitive to outliers and suffer from divergence for difficult cases. To address t…