Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

85169254338 · Jun 202019922001200920172026
48 results for mean distance

Ball k-means reduces point-centroid distance computations for faster k-means clustering.

problem Efficiently finding k-means clusters in large datasets.
method Uses a ball to describe clusters, dividing them into stable and active areas, and adjusting points within annulus areas.
result Significantly reduces point-centroid distance computations, making k-means faster and more efficient.

Study shows effective resistance distance yields more accurate network barycenter than Hamming distance.

problem Identifying the best metric for computing the Fréchet mean network.
method Compared the effectiveness of Hamming distance and effective resistance distance in capturing network topology.
result Effective resistance distance produces a more accurate Fréchet mean network.

We consider the setting of Reeb graphs of piecewise linear functions and study distances between them that are stable, meaning that functions which are similar in the supremum norm ought to have similar Reeb graphs. We define an edit distance for Reeb graphs and prove that it is stable and universal, meaning that it pr…

2018-01-05abs ↗pdf ↗

Study of metrics on positive-definite matrices from power potential, linking to power means.

problem Understanding metrics on positive-definite matrices derived from power potential.
method Explicit expressions for geodesics and distance function derived from Hessian of power potential.
result Geodesics and distance function converge to weighted matrix geometric mean as β tends to zero.

A new metric compares true and learned causal graphs considering data and graph structure.

problem Comparing true and learned causal graphs accurately.
method Continuous Structural Intervention Distance (CSID) using conditional mean embeddings and maximum mean discrepancy.
result Validated the CSID with synthetic data, showing its effectiveness in comparing causal graphs.

The original k-means clustering method works only if the exact vectors representing the data points are known. Therefore calculating the distances from the centroids needs vector operations, since the average of abstract data points is undefined. Existing algorithms can be extended for those cases when the sole input i…

2013-03-24abs ↗pdf ↗

Study robust mean estimation under coordinate-level corruptions using Hamming distance.

problem Robust mean estimation under realistic coordinate-level corruptions.
method Introduce a novel Hamming distance-based measure and present information-theoretic analysis.
result Data cleaning-inspired approaches can match information theoretic bounds for robust mean estimation.

In high dimensions, the mean and geometric median are nearly identical.

problem Understanding the relationship between mean and geometric median in high-dimensional spaces.
method Analytical derivation and simulation of the distance between mean and geometric median.
result The distance between mean and geometric median vanishes with dimensionality in high dimensions.

Paper presents a novel k-means clustering method using two distance measures for Gaussian data.

problem Improving clustering accuracy and robustness for Gaussian data.
method Integrates within cluster distance (WCD) and inter cluster distance (ICD) into k-means clustering.
result The algorithm provides more accurate clustering and better handles outliers.

A new Wasserstein KK-means method for clustering probability distributions.

problem Clustering probability distributions using the Wasserstein metric.
method Distance-based KK-means with SDP relaxation for Wasserstein barycenters.
result Distance-based KK-means outperforms centroid-based KK-means for clustering probability distributions.

Clustering is a separation of data into groups of similar objects. Every group called cluster consists of objects that are similar to one another and dissimilar to objects of other groups. In this paper, the K-Means algorithm is implemented by three distance functions and to identify the optimal distance function for c…

2013-03-11abs ↗pdf ↗

Two singular mean curvature flows converge if Hausdorff distance decreases faster than inverse time.

problem Understanding convergence of singular mean curvature flows.
method Analyzing Hausdorff distance between flows approaching a singularity.
result Two flows become identical if Hausdorff distance decreases faster than inverse time.

This work examines the sensitivity of energy distance to mean differences compared to covariance differences.

problem The sensitivity of energy distance to mean differences compared to covariance differences when distributions are close.
method Analyzes the energy distance in the case where distributions are close, focusing on sensitivity to mean and covariance differences.
result Energy distance is more sensitive to mean differences than covariance differences when distributions are close.

TS-K-means improves financial data clustering with dynamic time warping.

problem Inadequate handling of temporal dependencies in financial time series data.
method Integrates Dynamic Time Warping into Time Series K-means for financial data.
result TS-K-means outperforms traditional K-means in financial data analysis.

A new method using energy distance for ensemble and scenario reduction.

problem Solving complex dynamic and stochastic programs, especially in energy systems.
method Proposes a new method based on energy distance for ensemble and scenario reduction.
result Reduced scenario sets exhibit better statistical properties for energy distance than Wasserstein distance.

Due to the progressive growth of the amount of data available in a wide variety of scientific fields, it has become more difficult to ma- nipulate and analyze such information. Even though datasets have grown in size, the K-means algorithm remains as one of the most popular clustering methods, in spite of its dependenc…

2016-05-10abs ↗pdf ↗

We consider the problem of subspace estimation in a Bayesian setting. Since we are operating in the Grassmann manifold, the usual approach which consists of minimizing the mean square error (MSE) between the true subspace UU and its estimate U^\hat{U} may not be adequate as the MSE is not the natural metric in the Gra…

2011-01-18abs ↗pdf ↗

This work robustifies Wasserstein distance estimation with MoM estimators for outlier-polluted data.

problem Estimating Wasserstein distance between two distributions with outliers.
method Introducing MoM-based robust estimators for Wasserstein distance.
result Consistent MoM-based estimators for Wasserstein distance with convergence rates.

This paper refines MMD for domain adaptation by balancing intra-class and inter-class distances.

problem Balancing intra-class and inter-class distances for better feature discriminability in domain adaptation.
method The paper theoretically proves two facts about MMD and proposes a novel discriminative MMD method to balance intra-class and inter-class distances.
result The proposed method improves feature discriminability and outperforms state-of-the-art methods.

This thesis improves kernel-based distances for statistical inference and integration.

problem Efficiently measuring distances between probability distributions for robust and smooth modeling.
method Kernel-based distances, focusing on maximum mean discrepancy (MMD) and novel kernel quantile discrepancies.
result Improved MMD estimators for simulation-based inference and conditional expectations.

The L2L^2-metric or Fubini-Study metric on the non-linear Grassmannian of all submanifolds of type MM in a Riemannian manifold (N,g)(N,g) induces geodesic distance 0. We discuss another metric which involves the mean curvature and shows that its geodesic distance is a good topological metric. The vanishing phenomenon for…

2004-09-17abs ↗pdf ↗

Proposes a new method for posterior sampling using MMD with negative distance kernel.

problem Posterior sampling and conditional generative modeling.
method Approximates joint distribution using discrete Wasserstein gradient flows of MMD with negative distance kernel.
result Establishes an error bound for posterior distributions and proves the method is a Wasserstein gradient flow.

Mean embeddings provide an extremely flexible and powerful tool in machine learning and statistics to represent probability distributions and define a semi-metric (MMD, maximum mean discrepancy; also called N-distance or energy distance), with numerous successful applications. The representation is constructed as the e…

2018-02-13abs ↗pdf ↗

Constructs portfolios based on Hellinger distance to normal, finding market invariance.

problem Finding a market invariant for portfolio construction.
method Uses Hellinger distance to normal distribution for portfolio construction and analysis.
result Minimum Hellinger distance varies drastically between markets, suggesting market invariance.

The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.

problem Linear inverse problems with approximate priors.
method Maximum Entropy on the Mean (MEM) method with data-driven priors.
result Empirical mean convergence and estimates for prior differences based on epigraphical distance.

This paper presents a distance-based discriminative framework for learning with probability distributions. Instead of using kernel mean embeddings or generalized radial basis kernels, we introduce embeddings based on dissimilarity of distributions to some reference distributions denoted as templates. Our framework exte…

2018-03-01abs ↗pdf ↗

Learning algorithms for implicit generative models can optimize a variety of criteria that measure how the data distribution differs from the implicit model distribution, including the Wasserstein distance, the Energy distance, and the Maximum Mean Discrepancy criterion. A careful look at the geometries induced by thes…

2017-12-21abs ↗pdf ↗

Study proves inequality for hypersurfaces and shows almost extremals are close to Wulff shape.

problem Proving anisotropic extrinsic radius pinching inequality for hypersurfaces.
method Analyzes anisotropic mean curvatures and studies equality cases.
result Almost extremal hypersurfaces are close to Wulff shape.

New metrics improve quantum ensemble learning efficiency and power.

problem Quantum ensembles' distances poorly understood due to measurement constraints.
method Introduce MMD-kk hierarchy of integral probability metrics for quantum ensembles.
result MMD-kk requires fewer samples for full discriminative power at higher kk.

For a certain class of distributions, we prove that the linear programming relaxation of kk-medoids clustering---a variant of kk-means clustering where means are replaced by exemplars from within the dataset---distinguishes points drawn from nonoverlapping balls with high probability once the number of points drawn a…

2013-09-12abs ↗pdf ↗

Here, a non-linear analysis method is applied rather than classical one to study projective Finsler geometry. More intuitively, by means of an inequality on Ricci-Finsler curvature, a projectively invariant pseudo-distance is introduced and an analogous of Schwarz' lemma in Finsler geometry is proved. Next, the Schwarz…

2013-10-02abs ↗pdf ↗

Study shows distance to boundary is always attained on varifolds with bounded curvature.

problem Understanding varifolds with bounded mean curvature in Riemannian manifolds.
method Proves a barrier principle at infinity using sharp maximum principles.
result Distance to boundary is always attained on varifolds with bounded curvature.

New kernel improves MMDs with theoretical guarantees for gradient flows.

problem Non-smoothness of negative distance kernel in MMDs.
method Smoothed 1D absolute value function followed by fractional integral transform.
result Improved theoretical guarantees for Wasserstein gradient flows.