Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

59118177236 · Jun 202019922001200920172026
48 results for average distance

Improved average distance classifier for HDLSS settings with multiple population differences.

problem Poor performance of average distance classifier in HDLSS settings with location and scale differences.
method Proposed transformations to the average distance classifier to handle multiple population differences.
result The proposed classifiers perform well even when populations differ in other aspects than location and scale.

The paper introduces tests for high-dimensional independence using maximum and average distance correlations.

problem Testing independence in high-dimensional data.
method Characterizes consistency properties, compares test statistics, examines null distributions, and presents a fast chi-square-based procedure.
result The proposed tests are non-parametric and applicable to various metrics.

Study on predicting graph labels at nodes using local averaging and distance estimation.

problem Predicting graph labels at nodes given observations at other nodes.
method Local averaging and distance estimation methods for graph regression.
result Alternative methods can achieve standard nonparametric rates even when graph neighborhoods are too large or small.

Single particle reconstruction (SPR) from cryo-electron microscopy (EM) is a technique in which the 3D structure of a molecule needs to be determined from its contrast transfer function (CTF) affected, noisy 2D projection images taken at unknown viewing directions. One of the main challenges in cryo-EM is the typically…

2016-11-10abs ↗pdf ↗

In this paper we study the geometry of metric spheres in the curve complex of a surface, with the goal of determining the "average" distance between points on a given sphere. Averaging is not technically possible because metric spheres in the curve complex are countably infinite and do not support any invariant probabi…

2011-09-28abs ↗pdf ↗

Proposes MFSWB for marginal fairness in SWB, improving efficiency and performance.

problem Achieving marginal fairness in SWB averaging.
method Defining MFSWB as a constrained SWB problem, proposing two surrogate problems and a new slicing distribution.
result Surrogate MFSWB problems effectively minimize distances to marginals and encourage marginal fairness.

We prove that S^2 x S^2 satisfies an intermediate condition between having metrics with positive Ricci and positive sectional curvature. Namely, there exist metrics for which the average of the sectional curvatures of any two planes tangent at the same point, but separated by a minimum distance in the 2-Grassmannian, i…

2012-09-28abs ↗pdf ↗

New research shows SAA can outperform SA for Wasserstein barycenters.

problem Optimizing Wasserstein barycenters with entropy regularization.
method Comparison of Stochastic Approximation (SA) and Sample Average Approximation (SAA) for large-scale problems.
result SAA can be more efficient than SA for Wasserstein barycenters, especially in large-scale settings.

Paper improves Bayesian inference in federated learning with new algorithm VR-FALD*.

problem Bayesian inference in federated learning with communication bottlenecks and statistical heterogeneity.
method Federated Averaging Langevin Dynamics (FALD) and VR-FALD*.
result VR-FALD* corrects client drift due to statistical heterogeneity, improving convergence.

We consider the problem of allocating samples to a finite set of discrete distributions in order to learn them uniformly well in terms of four common distance measures: 22\ell_2^2, 1\ell_1, ff-divergence, and separation distance. To present a unified treatment of these distances, we first propose a general optimistic…

2019-10-28abs ↗pdf ↗

Paper estimates non-causal graphical models using covariance extension and transportation distance.

problem Estimating non-causal graphical models with smoothing relations.
method Proposes a covariance extension problem and uses transportation distance to minimize error with white noise.
result Solution is a double-sided autoregressive non-causal graphical model.

Paper proposes a new method for estimating treatment effects using interpretable deep learning models.

problem Estimating treatment effects from observational data with interpretability.
method Proposes a novel objective function using energy distance balancing score and neural additive models for improved interpretability.
result Demonstrates superior performance over state-of-the-art methods in semi-synthetic experiments.

A method for learning embeddings from multi-view data using Gromov-Wasserstein.

problem Challenges in learning low-dimensional representations from multi-view relational data with differing geometries.
method Bary-GWMDS and Mean-GWMDS-C, Gromov-Wasserstein-based methods operating on distance matrices.
result Stable and geometrically meaningful embeddings learned from synthetic and real-world datasets.

The original k-means clustering method works only if the exact vectors representing the data points are known. Therefore calculating the distances from the centroids needs vector operations, since the average of abstract data points is undefined. Existing algorithms can be extended for those cases when the sole input i…

2013-03-24abs ↗pdf ↗

Paper analyzes normal approximation for two-timescale stochastic algorithms, revealing interaction between fast and slow timescales.

problem Non-asymptotic bounds for accuracy of normal approximation in linear two-timescale stochastic approximation algorithms.
method Established bounds for normal approximation in terms of convex distance, focusing on last iterate and Polyak-Ruppert averaging.
result Normal approximation rate for the last iterate improves with increased timescale separation, while it decreases in the averaged setting.

The possibility that price dynamics is affected by its distance from a moving average has been recently introduced as new statistical tool. The purpose is to identify the tendency of the price dynamics to be attractive or repulsive with respect to its own moving average. We consider a number of tests for various models…

2006-01-12abs ↗pdf ↗

New method estimates SW distance using CDFs for scalable data parallelism.

problem Estimating SW distance efficiently for large datasets.
method Estimators based on CDFs of projected measures, avoiding sorting.
result Efficient estimation for large datasets and federated learning.

The paper studies parallel surfaces of cuspidal cross caps and their degeneracy.

problem Investigating the geometry and singularities of parallel surfaces of cuspidal cross caps.
method Established a criterion for the degeneracy of the distance squared function using geometric invariants.
result Parallel surfaces degenerate into a degenerated cuspidal S1 singularity at specific distances.

Unified meta algorithms estimate various distribution functionals in infinite-armed bandits.

problem Estimating various distribution functionals in infinite-armed bandits.
method Unified meta algorithms for offline and online settings, achieving optimal sample complexities.
result Online estimation offers significant advantage for certain distribution functionals.

New method improves sampling efficiency in complex stochastic systems.

problem Sampling efficiency in nonconvex stochastic gradient cases.
method Reflection coupling for unadjusted generalized Hamiltonian Monte Carlo.
result Quantitative Gaussian concentration bounds and convergence rates established.

Optimal transport (\OT) theory defines a powerful set of tools to compare probability distributions. \OT~suffers however from a few drawbacks, computational and statistical, which have encouraged the proposal of several regularized variants of OT in the recent literature, one of the most notable being the \textit{slice…

2019-02-01abs ↗pdf ↗

We define a C^1 distance between submanifolds of a riemannian manifold M and show that, if a compact submanifold N is not moved too much under the isometric action of a compact group G, there is a G-invariant submanifold C^1-close to N. The proof involves a procedure of averaging nearby submanifolds of riemannian manif…

1999-08-25abs ↗pdf ↗

Paper improves CLT and bootstrap approximations for LSA with decreasing step size.

problem Improving normal approximation and bootstrap methods for LSA with decreasing step sizes.
method Refined Berry-Esseen bounds and multiplier bootstrap procedure for LSA.
result Approximation rates up to 1/n1/\sqrt{n} for LSA rescaled error distribution.

This paper analyzes minibatch optimal transport distances and their applications.

problem Optimal transport distances are complex and impractical for large datasets.
method Extended analysis of minibatch optimal transport distances, focusing on various kernels and debiased functions.
result Minibatch optimal transport distances are unbiased estimators and have statistical and optimisation properties.

GANICE improves GAN-based causal inference by minimizing averaged Wasserstein risk.

problem Estimating interventional outcome distributions and quantiles in causal inference.
method GANICE uses extended Wasserstein distance and a cellwise critic to minimize averaged Wasserstein risk.
result GANICE achieves minimax optimality and consistently outperforms existing methods.

This work examines the sensitivity of energy distance to mean differences compared to covariance differences.

problem The sensitivity of energy distance to mean differences compared to covariance differences when distributions are close.
method Analyzes the energy distance in the case where distributions are close, focusing on sensitivity to mean and covariance differences.
result Energy distance is more sensitive to mean differences than covariance differences when distributions are close.

A new method using spherical harmonics approximates the Sliced-Wasserstein distance.

problem Approximating the Sliced-Wasserstein distance between probability measures.
method Spherical Harmonics Control Variates (SHCV) method for Monte Carlo approximation of the SW distance.
result SHCV method provides an improved rate of convergence compared to Monte Carlo for general measures.

Nearest Neighbors Algorithm is a Lazy Learning Algorithm, in which the algorithm tries to approximate the predictions with the help of similar existing vectors in the training dataset. The predictions made by the K-Nearest Neighbors algorithm is based on averaging the target values of the spatial neighbors. The selecti…

2018-11-13abs ↗pdf ↗

Testing two potentially multivariate variables for statistical dependence on the basis finite samples is a fundamental statistical challenge. Here we explore a family of tests that adapt to the complexity of the relationship between the variables, promising robust power across scenarios. Building on the distance correl…

2018-10-06abs ↗pdf ↗

Multi-label classification is a type of supervised learning where an instance may belong to multiple labels simultaneously. Predicting each label independently has been criticized for not exploiting any correlation between labels. In this paper we propose a novel approach, Nearest Labelset using Double Distances (NLDD)…

2017-02-15abs ↗pdf ↗

This work improves understanding of projection robust optimal transport distances.

problem Understanding the behavior of minimum Wasserstein estimators in high-dimensional and misspecified models.
method Adopting projection robust (PR) optimal transport, establishing statistical properties, proposing IPRW distance, and providing asymptotic guarantees.
result Established fundamental statistical properties and proposed new distances that outperform Wasserstein distances empirically.

Optimal posterior distributions improve SVM classifiers and parameter selection.

problem Improving SVM classifiers and selecting optimal regularization parameters.
method PAC-Bayesian approach with optimal posterior identification for stochastic classifiers.
result Optimal posteriors yield tight risk bounds and improved SVM performance.

This work improves scalability of Wasserstein distances in high dimensions.

problem Scalability issues in computing Wasserstein distances in high dimensions.
method Empirical convergence rates, robustness to data contamination, and computational methods.
result Established fast rates and robust estimation risks for sliced Wasserstein distances.