CLSVAE repairs systematic errors in images with minimal labeled data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We describe a formal approach to identify 'root causes' of outliers observed in variables in a scenario where the causal relation between the variables is a known directed acyclic graph (DAG). To this end, we first introduce a systematic way to define outlier scores. Further, we introduce the concep…
Survey compares methods for generating artificial outliers.
A new robust scaling approach improves downstream metabolomics analysis.
A new robust regression method handles outliers in high-dimensional data.
We perform an extended analysis of the distribution of drawdowns in the two leading exchange markets (US dollar against the Deutsmark and against the Yen), in the major world stock markets, in the U.S. and Japanese bond market and in the gold market, by introducing the concept of ``coarse-grained drawdowns,'' which all…
We propose a unified and systematic framework for performing online nonnegative matrix factorization in the presence of outliers. Our framework is particularly suited to large-scale data. We propose two solvers based on projected gradient descent and the alternating direction method of multipliers. We prove that the se…
Survey on reproducibility and distortion issues in text clustering and topic modeling.
Optimizing full likelihoods adapts loss scales and shapes for robust modeling.
Study improves robustness of Bayesian inference for cognitive models.
An Ensemble Anomaly Detection Framework for Risk Calculation Integrity
To identify and classify toxic online commentary, the modern tools of data science transform raw text into key features from which either thresholding or learning algorithms can make predictions for monitoring offensive conversations. We systematically evaluate 62 classifiers representing 19 major algorithmic families …
Robust statistics traditionally focuses on outliers, or perturbations in total variation distance. However, a dataset could be corrupted in many other ways, such as systematic measurement errors and missing covariates. We generalize the robust statistics approach to consider perturbations under any Wasserstein distance…
We propose an inlier-based outlier detection method capable of both identifying the outliers and explaining why they are outliers, by identifying the outlier-specific features. Specifically, we employ an inlier-based outlier detection criterion, which uses the ratio of inlier and test probability densities as a measure…
A novel approach ODAR detects outliers for clustering.
Clustering, or unsupervised classification, is a task often plagued by outliers. Yet there is a paucity of work on handling outliers in clustering. Outlier identification algorithms tend to fall into three broad categories: outlier inclusion, outlier trimming, and post hoc outlier identification methods, with the forme…
New framework detects outliers in non-IID categorical data.
Outlier detection aims to identify unusual data instances that deviate from expected patterns. The outlier detection is particularly challenging when outliers are context dependent and when they are defined by unusual combinations of multiple outcome variable values. In this paper, we develop and study a new conditiona…
Paper proposes methods to make OT robust to outliers.
Outlier detection plays an essential role in many data-driven applications to identify isolated instances that are different from the majority. While many statistical learning and data mining techniques have been used for developing more effective outlier detection algorithms, the interpretation of detected outliers do…
Advances in sensor technology have enabled the collection of large-scale datasets. Such datasets can be extremely noisy and often contain a significant amount of outliers that result from sensor malfunction or human operation faults. In order to utilize such data for real-world applications, it is critical to detect ou…
Paper tackles outlier detection in signals modeled by generative models with theoretical guarantees.
Transforms distance-based outlier scores into interpretable probabilistic estimates.
Generates synthetic data for benchmarking unsupervised outlier detection.
A novel unsupervised outlier detection method using Randomized PCA Forest.
HLoOP detects outliers in hyperbolic 2-space.
Robust PCA, the problem of PCA in the presence of outliers has been extensively investigated in the last few years. Here we focus on Robust PCA in the outlier model where each column of the data matrix is either an inlier or an outlier. Most of the existing methods for this model assumes either the knowledge of the dim…
New outlier detection method using graph Laplacian spectrum boosts performance.
We evaluate how modern outlier detection methods perform in identifying outliers in e-commerce conversion rate data. Based on the limitations identified, we then present a novel method to detect outliers in e-commerce conversion rate. This unsupervised method is made more business relevant by letting it automatically a…
Enhances clustering for functional data, robust to outliers.
A new method detects outliers using ensembles of Dirichlet process mixtures.
Rare data in a large-scale database are called outliers that reveal significant information in the real world. The subspace-based outlier detection is regarded as a feasible approach in very high dimensional space. However, the outliers found in subspaces are only part of the true outliers in high dimensional space, in…
Robust PCA, the problem of PCA in the presence of outliers has been extensively investigated in the last few years. Here we focus on Robust PCA in the column sparse outlier model. The existing methods for column sparse outlier model assumes either the knowledge of the dimension of the lower dimensional subspace or the …
ODIM detects outliers by under-fitting generative models, outperforming other methods.
A novel density-based approach QC detects outliers in data with high precision.
Inference in the presence of outliers is an important field of research as outliers are ubiquitous and may arise across a variety of problems and domains. Bayesian optimization is method that heavily relies on probabilistic inference. This allows outstanding sample efficiency because the probabilistic machinery provide…
Paper tackles outlier detection in multi-armed bandits, achieving high accuracy with reduced exploration costs.
We focus on the problem of unsupervised cell outlier detection and repair in mixed-type tabular data. Traditional methods are concerned only with detecting which rows in the dataset are outliers. However, identifying which cells are corrupted in a specific row is an important problem in practice, and the very first ste…
PyODDS automates outlier detection for new data sources.
This paper bridges outlier-robust estimation in robotics and computer vision with robust statistics.
A new method detects and corrects outliers using optimal transport.
Notwithstanding the popularity of conventional clustering algorithms such as K-means and probabilistic clustering, their clustering results are sensitive to the presence of outliers in the data. Even a few outliers can compromise the ability of these algorithms to identify meaningful hidden structures rendering their o…
This paper investigates differentially private analysis of distance-based outliers. The problem of outlier detection is to find a small number of instances that are apparently distant from the remaining instances. On the other hand, the objective of differential privacy is to conceal presence (or absence) of any partic…
Integrates outlier detection into neural networks for improved performance.
One-Class Boundary Peeling detects outliers efficiently and robustly.
Adaptive algorithm for outlier detection by balancing arm exploration and threshold estimation.
New method clusters matrix-variate data with outliers.
A new KF handles outliers without MSE loss.