This paper presents a general notion of Mahalanobis distance for functional data that extends the classical multivariate concept to situations where the observed data are points belonging to curves generated by a stochastic process. More precisely, a new semi-distance for functional observations that generalize the usu…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Extends Mahalanobis distance to Banach spaces for anomaly detection.
Flexible classifier using Mahalanobis distances for non-elliptical distributions.
The paper connects neural networks to Mahalanobis distance for interpretability.
A method uses Shapley values and Mahalanobis distances to explain multivariate outliers.
The Mahalanobis distance-based confidence score, a recently proposed anomaly detection method for pre-trained neural classifiers, achieves state-of-the-art performance on both out-of-distribution (OoD) and adversarial examples detection. This work analyzes why this method exhibits such strong performance in practical s…
Enhances forecasting of complex systems using FKMD.
Connects robust optimization to conformal prediction for uncertainty sets.
Bayesian nonparametric models improve OOD detection, especially with complex covariance structures.
Compressing data helps learn Mahalanobis metrics effectively.
The paper tackles learning smooth distance functions using query-based methods.
Proposes ridge regression on Riemannian manifolds for time-series prediction.
This paper learns user preferences from comparisons using Mahalanobis metrics.
A fundamental question in data analysis, machine learning and signal processing is how to compare between data points. The choice of the distance metric is specifically challenging for high-dimensional data sets, where the problem of meaningfulness is more prominent (e.g. the Euclidean distance between images). In this…
For many tasks and data types, there are natural transformations to which the data should be invariant or insensitive. For instance, in visual recognition, natural images should be insensitive to rotation and translation. This requirement and its implications have been important in many machine learning applications, a…
Paper proposes a robust metric learning algorithm.
The distance metric plays an important role in nearest neighbor (NN) classification. Usually the Euclidean distance metric is assumed or a Mahalanobis distance metric is optimized to improve the NN performance. In this paper, we study the problem of embedding arbitrary metric spaces into a Euclidean space with the goal…
A new test validates ensemble models against the null hypothesis.
Develops a method for stress testing correlations of financial portfolios.
This paper quantifies how hard it is to identify specific data points in machine learning models.
Extends metrics for SPD matrices to infinite dimensions.
Single particle reconstruction (SPR) from cryo-electron microscopy (EM) is a technique in which the 3D structure of a molecule needs to be determined from its contrast transfer function (CTF) affected, noisy 2D projection images taken at unknown viewing directions. One of the main challenges in cryo-EM is the typically…
Neural networks can learn distance metrics affecting model performance.
Recent work in metric learning has significantly improved the state-of-the-art in k-nearest neighbor classification. Support vector machines (SVM), particularly with RBF kernels, are amongst the most popular classification algorithms that uses distance metrics to compare examples. This paper provides an empirical analy…
The paper evaluates and improves uncertainty estimates in neural networks for safety-critical applications.
Boosts kernel two-sample test power with multiple kernels.
2L-FUSE enhances feature sparsity through kernel learning.
Elliptical Attention improves transformer performance by focusing on contextually relevant features.
Survey of spectral, probabilistic, and deep metric learning methods.
Kernel functions in support vector machines (SVM) are needed to assess the similarity of input samples in order to classify these samples, for instance. Besides standard kernels such as Gaussian (i.e., radial basis function, RBF) or polynomial kernels, there are also specific kernels tailored to consider structure in t…
Study quantifies distribution shifts and uncertainties to improve machine learning model robustness.
High-dimensional prediction is a challenging problem setting for traditional statistical models. Although regularization improves model performance in high dimensions, it does not sufficiently leverage knowledge on feature importances held by domain experts. As an alternative to standard regularization techniques, we p…
There is an increasingly apparent need for validating the classifications made by deep learning systems in safety-critical applications like autonomous vehicle systems. A number of recent papers have proposed methods for detecting anomalous image data that appear different from known inlier data samples, including reco…
This work briefly explores the possibility of approximating spatial distance (alternatively, similarity) between data points using the Isolation Forest method envisioned for outlier detection. The logic is similar to that of isolation: the more similar or closer two points are, the more random splits it will take to se…
This study examines the relationship between PLS and OLS regression using eigenvalue distributions.
Improved anomaly detection for launch vehicle propulsion systems using LSTM and statistical relabeling.
Study metric learning from limited preference comparisons, showing how low-dimensional structure can still reveal metric information.
Proposes GMOTE for better handling imbalanced data.
Tests if vertices in graphs have the same latent positions.
The objective of this study is to investigate the efficient determination of and for Support Vector Regression with RBF or mahalanobis kernel based on numerical and statistician considerations, which indicates the connection between and kernels and demonstrates that the deviation of geometric distance of ne…
Improved multivariate conformal prediction by standardizing residuals.
Metric learning is a key problem for many data mining and machine learning applications, and has long been dominated by Mahalanobis methods. Recent advances in nonlinear metric learning have demonstrated the potential power of non-Mahalanobis distance functions, particularly tree-based functions. We propose a novel non…
Unified AI detection framework for various artifacts.
Distance metric learning is a successful way to enhance the performance of the nearest neighbor classifier. In most cases, however, the distribution of data does not obey a regular form and may change in different parts of the feature space. Regarding that, this paper proposes a novel local distance metric learning met…
Nearest Neighbors Algorithm is a Lazy Learning Algorithm, in which the algorithm tries to approximate the predictions with the help of similar existing vectors in the training dataset. The predictions made by the K-Nearest Neighbors algorithm is based on averaging the target values of the spatial neighbors. The selecti…
Clustering and classification critically rely on distance metrics that provide meaningful comparisons between data points. We present mixed-integer optimization approaches to find optimal distance metrics that generalize the Mahalanobis metric extensively studied in the literature. Additionally, we generalize and impro…
This paper proposes the continuous semantic topic embedding model (CSTEM) which finds latent topic variables in documents using continuous semantic distance function between the topics and the words by means of the variational autoencoder(VAE). The semantic distance could be represented by any symmetric bell-shaped geo…
We consider an enlarged dimension reduction space in functional inverse regression. Our operator and functional analysis based approach facilitates a compact and rigorous formulation of the functional inverse regression problem. It also enables us to expand the possible space where the dimension reduction functions bel…