Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

69137206274 · Jun 202019922001200920172026
48 results for Directed distances

A new method optimizes projection directions for sliced Wasserstein distances.

problem Finding informative projecting directions for sliced Wasserstein distances is computationally expensive.
method Amortized projection optimization to predict directions efficiently.
result Proposed amortized models improve generative modeling performance.

The paper finds non-Gaussian directions in high-dimensional data using Wasserstein distance.

problem Locating interesting non-Gaussian features in high-dimensional data.
method Projection pursuit using 2-Wasserstein distance to maximize the difference from Gaussian.
result Statistical guarantees for accurately approximating an unknown low-dimensional non-Gaussian subspace.

We present a case-study demonstrating the usefulness of Bayesian hierarchical mixture modelling for investigating cognitive processes. In sentence comprehension, it is widely assumed that the distance between linguistic co-dependents affects the latency of dependency resolution: the longer the distance, the longer the …

2017-02-02abs ↗pdf ↗

In this paper, we deal with the problem of inferring causal directions when the data is on discrete domain. By considering the distribution of the cause P(X)P(X) and the conditional distribution mapping cause to effect P(YX)P(Y|X) as independent random variables, we propose to infer the causal direction via comparing the di…

2018-03-21abs ↗pdf ↗

Causal inference relies on the structure of a graph, often a directed acyclic graph (DAG). Different graphs may result in different causal inference statements and different intervention distributions. To quantify such differences, we propose a (pre-) distance between DAGs, the structural intervention distance (SID). T…

2013-06-05abs ↗pdf ↗

A new method optimizes slicing directions for SW distances to improve high-dimensional probability measure comparison.

problem Challenging identification of informative slicing directions for SW distances.
method Constrained learning approach to optimize slicing directions, using continuous relaxations and gradient-based primal-dual approach.
result Demonstrated efficacy in learning more informative slicing directions on various high-dimensional data.

A new distance metric for vMF distributions simplifies spherical data analysis.

problem Intractability of normalization constants and lack of suitable geometric metrics for comparing vMF distributions.
method Proposes a Wasserstein-like distance that decomposes vMF distribution discrepancies into angular and concentration components.
result The proposed distance metric induces a latent geometric structure on the space of non-degenerate vMF distributions.

Modified Wasserstein metric for Gaussian distributions, invariant to isometries.

problem Distance measurement for latent Gaussian distributions invariant to isometries.
method Modified Benamou-Brenier approach leading to a Procrustes Wasserstein metric.
result For Gaussian distributions, the metric reduces to Euclidean distance between eigenvalues.

TQF models multivariate uncertainty by learning conditional quantiles.

problem Challenges in fully nonparametric estimation of multivariate conditional distributions.
method Tomographic Quantile Forests (TQF) learns conditional quantiles of directional projections.
result TQF reconstructs multivariate conditional distribution efficiently without convexity restrictions.

Enhances LDL by integrating distance and directional information for more robust label feature representation.

problem Lack of robust label feature representation in LDL tasks, especially with label ambiguity.
method Introduces Structural Anchor Points (SAPs) to capture inter-cluster interactions and a novel LSFs construction strategy, LIFT-SAP.
result Improves LDL performance by 15% on average across 15 real-world datasets.

The distance function ϱ(p,q)\varrho(p,q) (or d(p,q)d(p,q)) of a distance space (general metric space) is not differentiable in general. We investigate such distance spaces over Rn\mathbb R^n, whose distance functions are differentiable like in case of Finsler spaces. These spaces have several good properties, yet they are no F…

2015-05-26abs ↗pdf ↗

This note demonstrates how both the concept of distance and the concept of holonomy can be constructed from a suitable network with directed edges (and no lengths). The number of different edge types depends on the signature of the metric and the dimension of the holonomy group. If the holonomy group is of dimension on…

2009-02-13abs ↗pdf ↗

Energy distance measures feature heterogeneity in federated learning.

problem Heterogeneity across data sources hinders model aggregation in federated learning.
method Introduced Taylor approximations of energy distance for efficient computation.
result Taylor approximations accurately capture feature discrepancies, improving convergence.

In the Engel group with its Carnot group structure we study subsets of locally finite subRiemannian perimeter and possessing constant subRiemannian normal. We prove the rectifiability of such sets: more precisely we show that, in some specific coordinates, they are upper-graphs of entire Lipschitz functions (with respe…

2012-01-30abs ↗pdf ↗

Sharp inequality between TV and Hellinger distances for Gaussian mixtures.

problem Understanding the relationship between total variation and Hellinger distances for Gaussian mixtures.
method Established a general upper bound on Hellinger distance in terms of TV distance raised to a power, demonstrating sharpness with specific examples.
result The Hellinger distance between two Gaussian mixtures is bounded by the TV distance raised to a power 1o(1)1-o(1), where o(1)o(1) is of order 1/loglog(1/TV)1/\log\log(1/\mathrm{TV}).

A new metric based on hitting probabilities for directed graphs and Markov chains.

problem Lack of metrics specifically adapted to asymmetric structure of directed graphs and Markov chains.
method Metric based on hitting probabilities, insensitive to shortest and average walk distances.
result New structural theory of directed graphs and utility for various applications.

Expands newsvendor model with moment constraints using Wasserstein distance.

problem Optimizing order quantity under distributional ambiguity.
method Formulates infinite dimensional primal problem, derives finite dimensional dual problem using problem of moments duality.
result Distributional ambiguity affects optimal order quantity and profits/costs.

We show that the recently introduced L1TV functional can be used to explicitly compute the flat norm for co-dimension one boundaries. While this observation alone is very useful, other important implications for image analysis and shape statistics include a method for denoising sets which are not boundaries or which ha…

2006-12-11abs ↗pdf ↗

This paper analyzes minibatch optimal transport distances and their applications.

problem Optimal transport distances are complex and impractical for large datasets.
method Extended analysis of minibatch optimal transport distances, focusing on various kernels and debiased functions.
result Minibatch optimal transport distances are unbiased estimators and have statistical and optimisation properties.

GANs excel at learning high dimensional distributions, but they can update generator parameters in directions that do not correspond to the steepest descent direction of the objective. Prominent examples of problematic update directions include those used in both Goodfellow's original GAN and the WGAN-GP. To formally d…

2018-02-13abs ↗pdf ↗

Autoencoders are a deep learning model for representation learning. When trained to minimize the distance between the data and its reconstruction, linear autoencoders (LAEs) learn the subspace spanned by the top principal directions but cannot learn the principal directions themselves. In this paper, we prove that $L_2…

2019-01-23abs ↗pdf ↗

This paper improves MDS visualization by adjusting Wasserstein distances for heavy-tailed data.

problem Enhancing Multidimensional Scaling (MDS) for better pattern recognition with heavy-tailed distributions.
method Introduces Max-D-SW, a metric adjustment of Max-Sliced Wasserstein distance that aggregates over orthonormal bases.
result Max-D-SW provides a clear numerical advantage in MDS outcomes, especially for heavy-tailed distributions.

We give a new proof of the Gromov theorem: For any C>0C>0 and integer n>1n>1 there exists a function ΔC,nΔ_{C,n} such that if the Gromov--Hausdorff distance between complete Riemannian nn-manifolds VV and WW is not greater than δδ, absolute values of their sectional curvatures KσC|K_σ|\leq C, and their injectivity radii…

2008-02-01abs ↗pdf ↗

Paper proposes PPMM for fast estimation of large-scale OTM.

problem Estimation of large-scale optimal transport maps (OTM) is challenging due to the curse of dimensionality.
method Combines projection pursuit regression and sufficient dimension reduction to adaptively select projection directions.
result PPMM consistently estimates the most informative projection direction and weakly converges to the target OTM.

MGDA converges under generalized smoothness for neural network optimization.

problem Optimizing neural networks with standard smoothness assumptions not holding.
method Revisited and analyzed MGDA and its stochastic version for generalized \ell-smooth MOO problems.
result MGDA and its variants converge to Pareto stationary points with guaranteed CA distance.

Self-supervised metric learning boosts downstream tasks in multi-view data.

problem Improving distance-based downstream tasks without labeled data.
method Developed a statistical framework to study self-supervised metric learning in multi-view data.
result Self-supervised metric learning improves target distances for various downstream tasks.

A fundamental question in data analysis, machine learning and signal processing is how to compare between data points. The choice of the distance metric is specifically challenging for high-dimensional data sets, where the problem of meaningfulness is more prominent (e.g. the Euclidean distance between images). In this…

2017-08-13abs ↗pdf ↗

The medoid of a set of n points is the point in the set that minimizes the sum of distances to other points. It can be determined exactly in O(n^2) time by computing the distances between all pairs of points. Previous works show that one can significantly reduce the number of distance computations needed by adaptively …

2019-06-11abs ↗pdf ↗