Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

6481,2971,9452,593 · Jun 202019922001200920172026
48 results for Concentration of distances

Paper addresses concentration of distances for fractional quasi p-norms, identifying conditions for concentration and anti-concentration.

problem Understanding concentration of distances for fractional quasi p-norms in high dimensions.
method Analyzes conditions for concentration and anti-concentration of distances for fractional quasi p-norms.
result Identifies conditions for concentration and anti-concentration of fractional quasi p-norms, ruling out some approaches and specifying conditions for control.

Study on volume of tubes and concentration in Riemannian geometry.

problem Understanding concentration loci in Riemannian manifolds and their relation to tube volumes.
method Provided a general formula for tube volumes, specialized to totally geodesic submanifolds, and investigated concentration loci.
result Explicitly proved concentration for codimension one cases and explored characterizations in Wasserstein and Box distances.

In this paper, we study the Lévy-Milman concentration phenomenon of 1-Lipschitz maps into infinite dimensional metric spaces. Our main theorem asserts that the concentration to an infinite dimensional p\ell^p-ball with the q\ell^q-distance function for 1p<q+1\leq p<q\leq +\infty is equivalent to the concentration to the…

2008-08-24abs ↗pdf ↗

A new method approximates the Sliced-Wasserstein distance without random projections.

problem Efficiently approximating the Sliced-Wasserstein distance for machine learning applications.
method Utilizing the concentration of measure phenomenon to develop a deterministic approximation.
result The approximation error goes to zero as the dimension increases, under a weak dependence condition.

Study on estimating distances between covariance operators and Gaussian processes.

problem Estimating distances between covariance operators and Gaussian processes.
method Riemannian distances, concentration results for Hilbert space-valued random variables, RKHS covariance and cross-covariance operators.
result Both distances converge in the Hilbert-Schmidt norm and can be consistently and efficiently estimated.

Improved estimation of concentration using half-spaces for adversarial vulnerability.

problem Understanding the concentration of measure phenomenon and its impact on adversarial vulnerability.
method Extending Gaussian Isoperimetric Inequality to non-spherical Gaussian measures and arbitrary ℓ_p-norms, using half-spaces to estimate concentration.
result Proposed method finds tighter intrinsic robustness bounds, providing evidence against concentration as a cause of adversarial vulnerability.

The paper analyzes how noise affects distances in high-dimensional data and when they remain useful.

problem Noise corrupts distances in high-dimensional data, making them unreliable for identifying true nearest and farthest neighbors.
method The paper uses asymptotic probabilistic expressions to characterize noise effects and decomposes data into ground truth and noise components.
result Under certain conditions, empirical neighborhood relations remain truthful even when distance concentration occurs.

New robust method for optimal transportation improves statistical inference.

problem Sensitivity to outliers and undefinedness in optimal transportation methods.
method Robust optimal transportation with a tuning parameter λ, leading to robust Wasserstein distance.
result The robust method provides statistical guarantees and improves machine learning applications.

This paper analyzes minibatch optimal transport distances and their applications.

problem Optimal transport distances are complex and impractical for large datasets.
method Extended analysis of minibatch optimal transport distances, focusing on various kernels and debiased functions.
result Minibatch optimal transport distances are unbiased estimators and have statistical and optimisation properties.

This paper shows how to estimate distances in latent space of random graphs using entropic OT.

problem Estimating distances between groups of nodes in latent space of random graphs.
method Entropic Optimal Transport (OT) with stability results for perturbations of the cost matrix.
result Consistent estimation of entropic OT distances between groups of nodes in latent space.

Constructs flows on manifolds with small curvature, proving Euclidean topology.

problem Geometric structure of manifolds with unbounded curvature.
method Distance like functions with integral hessian bound, Ricci flows.
result Manifolds with Ricci lower bound, non-negative scalar curvature, bounded entropy, Ahlfors nn-regular and small curvature concentration are topologically Euclidean.

A new distance metric for vMF distributions simplifies spherical data analysis.

problem Intractability of normalization constants and lack of suitable geometric metrics for comparing vMF distributions.
method Proposes a Wasserstein-like distance that decomposes vMF distribution discrepancies into angular and concentration components.
result The proposed distance metric induces a latent geometric structure on the space of non-degenerate vMF distributions.

We study the non-asymptotic behavior of a Coulomb gas on a compact Riemannian manifold. This gas is a symmetric n-particle Gibbs measure associated to the two-body interaction energy given by the Green function. We encode such a particle system by using an empirical measure. Our main result is a concentration inequalit…

2018-09-12abs ↗pdf ↗

We present a novel notion of outlier, called the Concentration Free Outlier Factor, or CFOF. As a main contribution, we formalize the notion of concentration of outlier scores and theoretically prove that CFOF does not concentrate in the Euclidean space for any arbitrary large dimensionality. To the best of our knowled…

2019-01-14abs ↗pdf ↗

Study on computing and estimating calibration distance, showing hardness and efficiency.

problem Computing and estimating calibration distance under different assumptions.
method Efficient algorithm for exact computation, polynomial-time approximation scheme; sample-based estimation for upper bounds.
result The problem becomes NP-hard when assumptions are removed, but efficient algorithms exist under certain conditions.

SGD converges to an invariant distribution with sub-Gaussian or sub-exponential properties.

problem Optimizing smooth and strongly convex objectives using SGD.
method Analysis through Markov chains, focusing on convergence and concentration properties.
result SGD iterates and their invariant limit distribution inherit sub-Gaussian or sub-exponential concentration properties.

Develops robust MDPs for unknown disturbances with performance guarantees.

problem Unknown disturbance distribution in MDPs.
method Empirical distribution, sublevel set of distance function, weak convergence, concentration inequality.
result Robust optimal value function converges to true optimal value function with increasing sample sizes.

The paper analyzes sparse high-dimensional linear regression with random design and unknown error variance, providing adaptiveness and concentration rates.

problem Sparse high-dimensional linear regression with random design and unknown error variance.
method Analysis of posterior concentration rates, employing techniques to address model misspecification.
result Adaptiveness and concentration rates of the posterior for sparse high-dimensional linear regression.

The study proves inequalities and curvature properties for Markov chains.

problem Isoperimetric and concentration inequalities for Markov chains.
method Laplacian separation principle for eikonal equation; modified log-Sobolev constant; Ollivier curvature.
result Affirmative answers to open questions and new inequalities.

We propose fast approximations for the generalized sliced-Wasserstein distance.

problem Efficient approximation of the generalized sliced-Wasserstein distance in high dimensions.
method Deterministic approximations using random projections and concentration of measure results.
result One-dimensional projections of high-dimensional random vectors are approximately Gaussian.

Develops a two-sample test using projected Wasserstein distance to handle high-dimensional data.

problem Testing whether two high-dimensional samples come from the same distribution.
method Optimal projection to find a low-dimensional linear mapping that maximizes the Wasserstein distance between projected probability distributions.
result Characterizes the convergence rate of the projected Wasserstein distance and presents practical algorithms.

New method improves sampling efficiency in complex stochastic systems.

problem Sampling efficiency in nonconvex stochastic gradient cases.
method Reflection coupling for unadjusted generalized Hamiltonian Monte Carlo.
result Quantitative Gaussian concentration bounds and convergence rates established.

The study uses heat flow to analyze properties of Laplace eigenfunctions on manifolds and domains.

problem Analyzing mass concentration and nodal domains of Laplace eigenfunctions.
method Heat diffusion technique to study eigenfunctions and their nodal sets.
result Discovers new insights into the decay and behavior of Laplace eigenfunctions.

Method detects effects of synthesis parameters on plutonium oxide microstructure.

problem Detecting effects of synthesis parameters on material microstructure.
method Copula theory, high dimensional distribution distances, and permutational statistics.
result Effects of strike order and oxalic acid feed on plutonium oxide microstructure detected.

Many statistical and machine learning approaches rely on pairwise distances between data points. The choice of distance metric has a fundamental impact on performance of these procedures, raising questions about how to appropriately calculate distances. When data points are real-valued vectors, by far the most common c…

2019-06-29abs ↗pdf ↗

We introduce an universum of the Polish (=complete separable metric) space - the convex cone of distance matrices and study its geometry. It happened that the generic Polish spaces in this sense of this universum is so called Urysohn spaces defined by P.S.Urysohn in 20-th, and generic metric triple (= metric space with…

2002-03-01abs ↗pdf ↗

Random feature matrices' singular values concentrate near their full expectation in high dimensions.

problem Characterizing the spectra of random feature matrices for regression problems.
method Analyzing two settings of input variables (random or well-separated) with conditions on dimension, complexity ratio, and sampling variance.
result The singular values of random feature matrices concentrate near their full expectation and near one with high probability.

Robustly clusters mixtures of Gaussians even with outliers.

problem Clustering mixtures of statistically separated Gaussians robustly to outliers.
method Uses certifiable hypercontractivity, bounded variance, and anti-concentration of linear projections.
result First efficient algorithm for robust clustering of statistically separated Gaussians mixtures.

For incomplete sub-Riemannian manifolds, and for an associated second-order hypoelliptic operator, which need not be symmetric, we identify two alternative conditions for the validity of Gaussian-type upper bounds on heat kernels and transition probabilities, with optimal constant in the exponent. Under similar conditi…

2018-10-15abs ↗pdf ↗

New approach to quantify posterior concentration rates using Wasserstein dynamics.

problem Quantifying the speed of posterior distribution concentration in Bayesian statistics.
method Combining local Lipschitz-continuity with dynamic formulation of Wasserstein distance.
result Optimal posterior contraction rates in finite and infinite-dimensional models.

Construction of ambiguity set in robust optimization relies on the choice of divergences between probability distributions. In distribution learning, choosing appropriate probability distributions based on observed data is critical for approximating the true distribution. To improve the performance of machine learning …

2017-05-23abs ↗pdf ↗

Adaptive sampling theory has shown that, with proper assumptions on the signal class, algorithms exist to reconstruct a signal in Rd\mathbb{R}^{d} with an optimal number of samples. We generalize this problem to the case of spatial signals, where the sampling cost is a function of both the number of samples taken and t…

2015-09-28abs ↗pdf ↗

Max-sliced Wasserstein metric reduces high-dimensional data to 1D for better estimation.

problem Curse of dimensionality in optimal transport.
method Introduces max-sliced Wasserstein metric to reduce high-dimensional problems to 1D.
result Uniform ratio bounds of empirical measures on RKHS concentrate uniformly fast at parametric rates.

This work aims to study the Portuguese regional agglomeration process, using the linear form the New Economic Geography models that emphasize the importance of spatial factors (distance, costs of transport and communication) in explaining of the concentration of economic activity in certain locations. In a theoretical …

2011-10-25abs ↗pdf ↗