A novel outlier score detects new road infrastructure images.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Transforms distance-based outlier scores into interpretable probabilistic estimates.
We study two problems in high-dimensional robust statistics: \emph{robust mean estimation} and \emph{outlier detection}. In robust mean estimation the goal is to estimate the mean of a distribution on given independent samples, an -fraction of which have been corrupted by a malicious…
Two new outlyingness scores improve outlier detection in high-dimensional data.
We describe a formal approach to identify 'root causes' of outliers observed in variables in a scenario where the causal relation between the variables is a known directed acyclic graph (DAG). To this end, we first introduce a systematic way to define outlier scores. Further, we introduce the concep…
A new score SiNNE improves OAM efficiency and accuracy.
Outlier detection amounts to finding data points that differ significantly from the norm. Classic outlier detection methods are largely designed for single data type such as continuous or discrete. However, real world data is increasingly heterogeneous, where a data point can have both discrete and continuous attribute…
Outlier detection aims to identify unusual data instances that deviate from expected patterns. The outlier detection is particularly challenging when outliers are context dependent and when they are defined by unusual combinations of multiple outcome variable values. In this paper, we develop and study a new conditiona…
A new framework detects adverse dataset shifts using outlier scores.
A new robust PCA method uses Innovation Search and Leverage Scores.
EHBOS enhances HBOS by capturing feature interactions, improving anomaly detection.
DTOR explains anomalies with rule-based explanations.
We present a novel notion of outlier, called the Concentration Free Outlier Factor, or CFOF. As a main contribution, we formalize the notion of concentration of outlier scores and theoretically prove that CFOF does not concentrate in the Euclidean space for any arbitrary large dimensionality. To the best of our knowled…
Selecting and combining the outlier scores of different base detectors used within outlier ensembles can be quite challenging in the absence of ground truth. In this paper, an unsupervised outlier detector combination framework called DCSO is proposed, demonstrated and assessed for the dynamic selection of most compete…
A novel unsupervised outlier detection method using Randomized PCA Forest.
High-dimensional data poses unique challenges in outlier detection process. Most of the existing algorithms fail to properly address the issues stemming from a large number of features. In particular, outlier detection algorithms perform poorly on data set of small size with a large number of features. In this paper, w…
SDCOR clusters massive datasets efficiently, detecting outliers with low memory usage.
Extends importance sampling to nonlinear models using adjoint operators.
FUSE neural centrality framework improves data point measurement in high dimensions.
Paper proposes a scoring function for detecting anomalies in large datasets.
Proposes GMOTE for better handling imbalanced data.
We focus on the problem of unsupervised cell outlier detection and repair in mixed-type tabular data. Traditional methods are concerned only with detecting which rows in the dataset are outliers. However, identifying which cells are corrupted in a specific row is an important problem in practice, and the very first ste…
Outlier detection plays an essential role in many data-driven applications to identify isolated instances that are different from the majority. While many statistical learning and data mining techniques have been used for developing more effective outlier detection algorithms, the interpretation of detected outliers do…
Two new scoring methods improve anomaly detection in Isolation Forest.
HLoOP detects outliers in hyperbolic 2-space.
Proposes MVG-CRPS for robust multivariate forecasting.
Rare data in a large-scale database are called outliers that reveal significant information in the real world. The subspace-based outlier detection is regarded as a feasible approach in very high dimensional space. However, the outliers found in subspaces are only part of the true outliers in high dimensional space, in…
Improved isolation forest for better outlier detection.
This paper presents a simple but effective density-based outlier detection approach with the local kernel density estimation (KDE). A Relative Density-based Outlier Score (RDOS) is introduced to measure the local outlierness of objects, in which the density distribution at the location of an object is estimated with a …
New framework detects outliers in non-IID categorical data.
Feature selection places an important role in improving the performance of outlier detection, especially for noisy data. Existing methods usually perform feature selection and outlier scoring separately, which would select feature subsets that may not optimally serve for outlier detection, leading to unsatisfying perfo…
MMDCP improves outlier detection and classification with adaptive prediction sets.
Integrates outlier detection into neural networks for improved performance.
Enhances predictive models against misspecification and outliers.
A new method detects outliers in dirty data using a leave-out strategy.
Outlier detection (also known as anomaly detection or deviation detection) is a process of detecting data points in which their patterns deviate significantly from others. It is common to have outliers in industry applications, which could be generated by different causes such as human error, fraudulent activities, or …
UN-AVOIDS visualizes and detects anomalies without needing labeled data.
Often the challenge associated with tasks like fraud and spam detection[1] is the lack of all likely patterns needed to train suitable supervised learning models. In order to overcome this limitation, such tasks are attempted as outlier or anomaly detection tasks. We also hypothesize that out- liers have behavioral pat…
NLR models often perform worse than LR for outlying input data in environmental sciences.
Geometric framework detects outliers in high-dimensional data.
Proposes ATH for KPI anomaly detection based on local data properties.
This paper investigates differentially private analysis of distance-based outliers. The problem of outlier detection is to find a small number of instances that are apparently distant from the remaining instances. On the other hand, the objective of differential privacy is to conceal presence (or absence) of any partic…
A novel framework IMBoost improves outlier detection by leveraging the inlier memorization effect.
ECOD detects outliers without parameters, fast and simple.
Study improves data quality assessment for structural monitoring data.
Outlier detection methods have become increasingly relevant in recent years due to increased security concerns and because of its vast application to different fields. Recently, Pauwels and Lasserre (2016) noticed that the sublevel sets of the inverse Christoffel function accurately depict the shape of a cloud of data …
Outlier detection is a technique in data mining that aims to detect unusual or unexpected records in the dataset. Existing outlier detection algorithms have different pros and cons and exhibit different sensitivity to noisy data such as extreme values. In this paper, we propose a novel cluster-based outlier detection a…
Robust score matching improves parameter estimation in contaminated data.