GCF estimates heterogeneous treatment effects for continuous treatments in online marketplaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Transforms distance-based outlier scores into interpretable probabilistic estimates.
Paper explains distance-based classifiers using neural network structures.
Neural networks learn distance-based representations, not just intensity.
We consider classifiers for high-dimensional data under the strongly spiked eigenvalue (SSE) model. We first show that high-dimensional data often have the SSE model. We consider a distance-based classifier using eigenstructures for the SSE model. We apply the noise reduction methodology to estimation of the eigenvalue…
The paper improves uncertainty quantification for node classification using distance-based regularization.
Solves clustering contradictions by high-dimensional embedding with wide gaps.
BRDAD uses bagging and regularization to improve anomaly detection without labeled data.
Improves classification of microbiome data using mixture distributions.
A new measure identifies clusters without assuming data distribution.
Time series classification is an increasing research topic due to the vast amount of time series data that are being created over a wide variety of fields. The particularity of the data makes it a challenging task and different approaches have been taken, including the distance based approach. 1-NN has been a widely us…
A method for ranking items using distance-based learning from positive and unlabeled data.
Extends conformal prediction to contrastive learning for better coverage of positive samples.
Anomalies (unusual patterns) in time-series data give essential, and often actionable information in critical situations. Examples can be found in such fields as healthcare, intrusion detection, finance, security and flight safety. In this paper we propose new conformalized density- and distance-based anomaly detection…
Extends RF proximities to all supervised distance-based machine learning contexts.
For time series comparisons, it has often been observed that z-score normalized Euclidean distances far outperform the unnormalized variant. In this paper we show that a z-score normalized, squared Euclidean Distance is, in fact, equal to a distance based on Pearson Correlation. This has profound impact on many distanc…
DBLE improves confidence calibration of DNNs by learning distances in representation space.
Neural networks can learn distance metrics affecting model performance.
GICDM corrects hubness in embedding spaces for better generative model evaluation.
ClusTR improves clustering-based models' robustness without adversarial training.
A new Wasserstein -means method for clustering probability distributions.
In general, the clustering problem is NP-hard, and global optimality cannot be established for non-trivial instances. For high-dimensional data, distance-based methods for clustering or classification face an additional difficulty, the unreliability of distances in very high-dimensional spaces. We propose a distance-ba…
Learning expressive low-dimensional representations of ultrahigh-dimensional data, e.g., data with thousands/millions of features, has been a major way to enable learning methods to address the curse of dimensionality. However, existing unsupervised representation learning methods mainly focus on preserving the data re…
In light of the power problems of statistical tests and undisciplined use of alpha-based statistics to compare models, this paper proposes a unified set of distance-based performance metrics, derived as the square root of the sum of squared alphas and squared standard errors. The Bayesian investor views model performan…
New metric learning approach for tree data reduces computation cost.
Despite numerous attempts to defend deep learning based image classifiers, they remain susceptible to the adversarial attacks. This paper proposes a technique to identify susceptible classes, those classes that are more easily subverted. To identify the susceptible classes we use distance-based measures and apply them …
A new robust time series distance metric for k-NN classification.
Distance plays a fundamental role in measuring similarity between objects. Various visualization techniques and learning tasks in statistics and machine learning such as shape matching, classification, dimension reduction and clustering often rely on some distance or similarity measure. It is of tremendous importance t…
The reliable measurement of confidence in classifiers' predictions is very important for many applications and is, therefore, an important part of classifier design. Yet, although deep learning has received tremendous attention in recent years, not much progress has been made in quantifying the prediction confidence of…
Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample and independence tests from the machine learning community. A fixed-point transf…
New GP kernel handles mixed-categorical data, improving model accuracy.
Study evaluates initialization strategies for infinite hidden Markov models.
Outliers are ubiquitous in modern data sets. Distance-based techniques are a popular non-parametric approach to outlier detection as they require no prior assumptions on the data generating distribution and are simple to implement. Scaling these techniques to massive data sets without sacrificing accuracy is a challeng…
The -nearest neighbour (-NN) classifier is one of the oldest and most important supervised learning algorithms for classifying datasets. Traditionally the Euclidean norm is used as the distance for the -NN classifier. In this thesis we investigate the use of alternative distances for the -NN classifier. We …
This paper explores ratio-based loss functions for machine learning.
Bayesian hierarchical clustering (BHC) is an agglomerative clustering method, where a probabilistic model is defined and its marginal likelihoods are evaluated to decide which clusters to merge. While BHC provides a few advantages over traditional distance-based agglomerative clustering algorithms, successive evaluatio…
Unified score and distance-based GoF tests for model adequacy.
Paper proposes a new method to quantify uncertainty in machine learning models.
A statistical model predicts generalization in few-shot learning.
Efficient method classifies locally stationary time series based on second-order characteristics.
Mahalanobis distance detects anomalies well, but not for classification.
In this paper, we define a reduced distance function based at a point at the singular time of a Ricci flow. We also show the monotonicity of the corresponding reduced volume based at time T, with equality iff the Ricci flow is a gradient shrinking soliton. Our curvature bound assumption is more general than …
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…
It is a key to construct a similarity graph in graph-oriented subspace learning and clustering. In a similarity graph, each vertex denotes a data point and the edge weight represents the similarity between two points. There are two popular schemes to construct a similarity graph, i.e., pairwise distance based scheme an…
Improves uncertainty estimation and OOD detection in neural networks.
Random forest can be adapted for open-set recognition with improved performance.
The paper compares clustering methods for improving time series forecasting accuracy.
This paper introduces the concept of kernels on fuzzy sets as a similarity measure for -valued functions, a.k.a. \emph{membership functions of fuzzy sets}. We defined the following classes of kernels: the cross product, the intersection, the non-singleton and the distance-based kernels on fuzzy sets. Applicabili…