Outlier detection is a fundamental task in data mining and has many applications including detecting errors in databases. While there has been extensive prior work on methods for outlier detection, modern datasets often have sizes that are beyond the ability of commonly used methods to process the data within a reasona…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes methods to make OT robust to outliers.
A large dimensional characterization of robust M-estimators of covariance (or scatter) is provided under the assumption that the dataset comprises independent (essentially Gaussian) legitimate samples as well as arbitrary deterministic samples, referred to as outliers. Building upon recent random matrix advances in the…
PRAE identifies outliers and reconstructs inliers in autoencoders.
Efficiently estimates sparse linear regression with heavy-tailed and outlier-contaminated data.
Identifies arms with rewards significantly different from the majority.
Adaptive algorithm for outlier detection by balancing arm exploration and threshold estimation.
A CAE improves DNN's outlier and adversary defense.
Clustering, or unsupervised classification, is a task often plagued by outliers. Yet there is a paucity of work on handling outliers in clustering. Outlier identification algorithms tend to fall into three broad categories: outlier inclusion, outlier trimming, and post hoc outlier identification methods, with the forme…
Often the challenge associated with tasks like fraud and spam detection[1] is the lack of all likely patterns needed to train suitable supervised learning models. In order to overcome this limitation, such tasks are attempted as outlier or anomaly detection tasks. We also hypothesize that out- liers have behavioral pat…
Study reveals significant performance flips in GLOD using repurposed graph classification datasets.
Proposes CE-BASS for robust Kalman filtering with innovative and additive outliers.
Study improves robustness and sparsity in linear regression with adversarial outliers and heavy-tailed noise.
We introduce and develop a novel approach to outlier detection based on adaptation of random subspace learning. Our proposed method handles both high-dimension low-sample size and traditional low-dimensional high-sample size datasets. Essentially, we avoid the computational bottleneck of techniques like minimum covaria…
In this work we perform outlier detection using ensembles of neural networks obtained by variational approximation of the posterior in a Bayesian neural network setting. The variational parameters are obtained by sampling from the true posterior by gradient descent. We show our outlier detection results are comparable …
A novel framework IMBoost improves outlier detection by leveraging the inlier memorization effect.
Few random images can improve anomaly detection performance.
Transforms distance-based outlier scores into interpretable probabilistic estimates.
We propose a novel procedure for outlier detection in functional data, in a semi-supervised framework. As the data is functional, we consider the coefficients obtained after projecting the observations onto orthonormal bases (wavelet, PCA). A multiple testing procedure based on the two-sample test is defined in order t…
Paper introduces detect-then-impute conformal prediction for cellwise outliers.
Inference in the presence of outliers is an important field of research as outliers are ubiquitous and may arise across a variety of problems and domains. Bayesian optimization is method that heavily relies on probabilistic inference. This allows outstanding sample efficiency because the probabilistic machinery provide…
A new framework detects adverse dataset shifts using outlier scores.
This paper proposes a new RWO-Sampling (Random Walk Over-Sampling) based on graphs for imbalanced datasets. In this method, two schemes based on under-sampling and over-sampling methods are introduced to keep the proximity information robust to noises and outliers. After constructing the first graph on minority class, …
New SMC sampler improves diffusion model sampling efficiency.
Paper develops a method to robustly cluster tensors with outliers.
New method falsifies causal graphs using outlier events.
SPCA improves PCA by learning from simple to complex samples.
Improved spectroscopy classification with deep learning and synthetic data.
MOSAIC selects few informative exemplars from high-dimensional data with non-linear structures.
Paper proposes a matrix optimization model for reliable Euclidean embedding from noisy data.
The paper develops p-values for outlier detection using conformal inference.
Online TERM improves robustness and fairness in streaming data.
AutoOD automates outlier detection using curiosity-guided search and self-imitation learning.
Integrates outlier detection into neural networks for improved performance.
We study two problems in high-dimensional robust statistics: \emph{robust mean estimation} and \emph{outlier detection}. In robust mean estimation the goal is to estimate the mean of a distribution on given independent samples, an -fraction of which have been corrupted by a malicious…
Efficiently estimates sparse linear regression with heavy-tailed data and outliers.
We consider the multi-class classification problem when the training data and the out-of-sample test data may have different distributions and propose a method called BCOPS (balanced and conformal optimized prediction sets). BCOPS constructs a prediction set as a subset of class labels, possibly empty. It tries …
Feature selection places an important role in improving the performance of outlier detection, especially for noisy data. Existing methods usually perform feature selection and outlier scoring separately, which would select feature subsets that may not optimally serve for outlier detection, leading to unsatisfying perfo…
WPCA improves subspace recovery robustness to outliers.
Enhances predictive models against misspecification and outliers.
This paper examines the problem of locating outlier columns in a large, otherwise low-rank matrix, in settings where {}{the data} are noisy, or where the overall matrix has missing elements. We propose a randomized two-step inference framework, and establish sufficient conditions on the required sample complexities und…
Paper proposes a method to generate synthetic anomalies for robust anomaly detection.
MFRDE uses medians of forest estimators to robustly estimate densities in noisy data.
Outlier detection is a crucial part of robust evaluation for crowdsourceable assessment of Quality of Experience (QoE) and has attracted much attention in recent years. In this paper, we propose some simple and fast algorithms for outlier detection and robust QoE evaluation based on the nonconvex optimization principle…
robROSE tackles imbalanced fraud data by creating synthetic samples and detecting outliers.
Linear regression models contaminated by Gaussian noise (inlier) and possibly unbounded sparse outliers are common in many signal processing applications. Sparse recovery inspired robust regression (SRIRR) techniques are shown to deliver high quality estimation performance in such regression models. Unfortunately, most…
Outlier detection is an important topic in machine learning and has been used in a wide range of applications. In this paper, we approach outlier detection as a binary-classification issue by sampling potential outliers from a uniform reference distribution. However, due to the sparsity of data in high-dimensional spac…
This paper examines the problem of locating outlier columns in a large, otherwise low-rank, matrix. We propose a simple two-step adaptive sensing and inference approach and establish theoretical guarantees for its performance; our results show that accurate outlier identification is achievable using very few linear sum…