A new method learns meaningful distances between samples using optimal transport.
problem Learning meaningful distances between samples in datasets without labeled data.
method Computes OT distances between samples and features using singular vectors of a function mapping ground metrics to OT distances.
result Wasserstein Singular Vectors provide a scalable solution for unsupervised ground metric learning.
Novel distances between distributions using conditional ground distances.
problem Quantifying distances between statistical multivariate distributions.
method Optimal transport with entropic regularization and ground distance on conditionals.
result Upper bounds for jointly convex distances and improved GMM learning.
Transportation distances have been used for more than a decade now in machine learning to compare histograms of features. They have one parameter: the ground metric, which can be any metric between the features themselves. As is the case for all parameterized distances, transportation distances can only prove useful in…
Unsupervised clustering can reproduce categorization systems if features and metrics are correctly selected.
problem Reproducing expert-provided categorization systems using unsupervised clustering.
method Investigated using toy datasets and real-world fund categorization. Used appropriate feature selection and a supervised Random Forest-based distance metric.
result Unsupervised clustering can reproduce ground truth classes if features and metrics are correctly selected.
Optimizes ground metric on graphs for evolving density models.
problem Optimizing ground metric for evolving density models.
method Adaptive ground metric learning constrained to geodesic distances on graphs.
result Efficiently learned geodesic distances align with observed density evolution.
Parts of Texas, Oklahoma, and Kansas have experienced increased rates of seismicity in recent years, providing new datasets of earthquake recordings to develop ground motion prediction models for this particular region of the Central and Eastern North America (CENA). This paper outlines a framework for using Artificial…
New causal distances improve evaluation of causal discovery algorithms.
problem Evaluating causal discovery algorithms using graphical distances is limited.
method Defined causal distances based on causal distributions rather than graphical structure.
result Improved evaluation of causal discovery algorithms on synthetic and real-world datasets.
The paper analyzes how noise affects distances in high-dimensional data and when they remain useful.
problem Noise corrupts distances in high-dimensional data, making them unreliable for identifying true nearest and farthest neighbors.
method The paper uses asymptotic probabilistic expressions to characterize noise effects and decomposes data into ground truth and noise components.
result Under certain conditions, empirical neighborhood relations remain truthful even when distance concentration occurs.
Wasserstein GANs use different p-metrics to improve model performance.
problem Improving stability and performance of Wasserstein GANs.
method Introduce (q,p)-Wasserstein GANs using various p-metrics. result Different p-metrics can notably improve GAN performance. Paper introduces metrics to evaluate missing data imputation without ground truth.
problem Handling missing data in time series without ground truth.
method Introduces Wasserstein distance (WD) and Jensen-Shannon divergence (JSD) as metrics to evaluate imputation quality.
result WD and JSD are effective metrics for assessing missing data imputation quality.
DBLE improves confidence calibration of DNNs by learning distances in representation space.
problem Poor confidence calibration of deep neural networks (DNNs).
method DBLE trains a confidence model jointly with the classification model, using distances in the representation space.
result DBLE outperforms alternative single-model confidence calibration approaches and ensemble methods.
GCNs' performance linked to feature, graph, and ground truth alignment.
problem Improving GCNs' classification performance.
method Subspace alignment measure (SAM) based on Frobenius norm of chordal distances.
result SAM quantifies the alignment between features, graph, and ground truth.
Study heat profiles and eigenfunctions using Brownian motion.
problem Investigate heat profiles and eigenfunctions of Laplace equations.
method Probabilistic tools based on Brownian motion and Feynman-Kac formulae.
result Supremum norm bounds for ground state Dirichlet eigenfunctions and comparison of maximum temperatures.
Bayesian distance clustering improves robustness to kernel choice.
problem Kernel sensitivity in model-based clustering.
method Modeling pairwise distances instead of original data.
result Dramatic gains in cluster inference robustness.
Proposes a new method for posterior sampling using MMD with negative distance kernel.
problem Posterior sampling and conditional generative modeling.
method Approximates joint distribution using discrete Wasserstein gradient flows of MMD with negative distance kernel.
result Establishes an error bound for posterior distributions and proves the method is a Wasserstein gradient flow.
Paper introduces untangling number to quantify 3-periodic tangle complexity.
problem Quantifying the complexity of 3-periodic tangles in biological, chemical, and physical systems.
method Introduces untangling number, a measure of minimum distance to ground state through diagrammatic operations.
result For infinite open curves, generic ground states are crystallographic rod packings.
New PAC-Bayesian bounds improve Sliced-Wasserstein distances.
problem Improving statistical properties of Sliced-Wasserstein distances.
method Leveraging PAC-Bayesian theory to provide bounds and learning procedures.
result PAC-Bayesian generalization bounds for adaptive SW distances.
The paper analyzes conditions for solving low-rank matrix recovery problems with noisy measurements.
problem Low-rank matrix recovery with corrupted measurements.
method Analysis of the restricted isometry property (RIP) and local search methods.
result Sharp bounds on the maximum distance between local minimizers and the ground truth.
Paper defines untangling number to measure entanglement complexity in 3-periodic networks.
problem Measuring the complexity of entanglement in 3-periodic networks.
method Defining ground states through knot-theoretic crossing diagrams and measuring untangling number.
result Introduced untangling number as a measure of entanglement complexity.
A new method efficiently approximates Gromov-Wasserstein distance.
problem High computational complexity of Gromov-Wasserstein distance.
method Importance sparsification method to construct a sparse coupling matrix.
result Efficient approximation of GW distance with reduced complexity.
Distance-based hierarchical clustering (HC) methods are widely used in unsupervised data analysis but few authors take account of uncertainty in the distance data. We incorporate a statistical model of the uncertainty through corruption or noise in the pairwise distances and investigate the problem of estimating the HC…
Images obtained with coherent illumination, as is the case of sonar, ultrasound-B, laser and Synthetic Aperture Radar -- SAR, are affected by speckle noise which reduces the ability to extract information from the data. Specialized techniques are required to deal with such imagery, which has been modeled by the G0 dist…
A new tree-sliced Wasserstein distance improves optimal transport computations.
problem Computational and statistical drawbacks in optimal transport.
method Introducing tree metrics and averaging Wasserstein distances using random tree metrics.
result Tree-sliced Wasserstein distance outperforms other methods on benchmarks.
New method estimates covariance in deep heteroscedastic regression without labels.
problem Estimating covariance in deep heteroscedastic models is challenging due to sample-dependent covariance and lack of ground truth.
method Proposes a self-supervised approach using KL Divergence and 2-Wasserstein distance for covariance estimation and a neighborhood-based heuristic for pseudo labels.
result Demonstrates effective pseudo labels and a computationally cheaper yet accurate deep heteroscedastic regression.
Recent work in distance metric learning has focused on learning transformations of data that best align with provided sets of pairwise similarity and dissimilarity constraints. The learned transformations lead to improved retrieval, classification, and clustering algorithms due to the better adapted distance or similar…
A clustering algorithm based on the Hausdorff distance is introduced and compared to the single and complete linkage. The three clustering procedures are applied to a toy example and to the time series of financial data. The dendrograms are scrutinized and their features confronted. The Hausdorff linkage relies of firm…
Method learns relational features for Gaifman models from knowledge bases.
problem Structure learning for Gaifman models.
method Relational tree distances to learn relational features.
result Empirical evaluation shows superiority over classical rule-learning.
Proposes HOT method for robust multi-view learning.
problem Inability of traditional methods to handle unaligned and non-distributionally aligned views.
method Hierarchical optimal transport (HOT) method that penalizes sliced Wasserstein distances between different views.
result HOT method achieves robust performance on both synthetic and real-world tasks.
Paper introduces CWDAE for better synthetic data generation.
problem Measuring discrepancy between generative and ground-truth distributions.
method Introduces mixture Cramer-Wold distance for joint and marginal distributional learning.
result CWDAE shows remarkable performance in generating synthetic data.
Generative model synthesizes earthquake acceleration data.
problem Robust estimation of ground motions for engineering applications.
method Wasserstein GAN formulation for conditioning on physical variables.
result Trained model synthesizes realistic 3-component accelerograms.
We consider a bounded domain Ω of RN, N≥3, and h a continuous function on Ω. Let Γ be a closed curve contained in Ω. We study existence of positive solutions u∈H01(Ω) to the equation −Δu+hu=ρΓ−σu2σ∗−1 in Ω where 2σ∗:=N−22(N−σ), $σ\in (0,2)…
Improved anomaly detection for launch vehicle propulsion systems using LSTM and statistical relabeling.
problem Detecting anomalies in real-time telemetry data for launch vehicles.
method Utilized LSTM networks for anomaly classification, introduced a statistical detector based on Mahalanobis distance and forward-backward detection fractions to adjust training labels.
result Precision and recall of the LSTM classifier improved by 7% and 22% respectively after statistical relabeling.
This paper establishes the consistency of spectral approaches to data clustering. We consider clustering of point clouds obtained as samples of a ground-truth measure. A graph representing the point cloud is obtained by assigning weights to edges based on the distance between the points they connect. We investigate the…
Model tracks structural changes in Brownian particle configurations on a sphere.
problem Tracking structural changes in Brownian particle configurations on a sphere.
method Introduces Frustrated Distance Matrix (FDM) model for dynamic distance matrices on S^2.
result Preserves static BBS template with dynamics as redistributed spectral mass.
Proposes TNCM-VAE for generating causal financial time series.
problem Lack of causal reasoning in market generators.
method Combines VAE with structural causal models, enforcing causal constraints through DAGs and using causal Wasserstein distance.
result Superior performance in counterfactual probability estimation, L1 distances as low as 0.03-0.10.
In this work, we present a method to compute the Kantorovich-Wasserstein distance of order one between a pair of two-dimensional histograms. Recent works in Computer Vision and Machine Learning have shown the benefits of measuring Wasserstein distances of order one between histograms with n bins, by solving a classic…
New model accounts for continuous human trajectories in robotics.
problem Inaccurate probabilistic models of human behavior in robotics.
method Developed a new probabilistic model that considers distances between continuous trajectories.
result The new model outperforms existing models in explaining human behavior and improving robot inference.
A new method matches measures across different spaces using cost-regularized optimal transport.
problem Matching measures in different spaces without aligned data.
method Cost-regularized optimal transport formulation to match measures across two Euclidean spaces.
result Demonstrated applicability to single-cell spatial transcriptomics/multiomics matching tasks.
Study improves estimation of functions from noisy data using convex penalties.
problem Estimating functions from noisy point evaluations of linear operators.
method Tikhonov regularization with convex and p-homogeneous penalty functionals. result Derives concentration rates for regularized solutions in symmetric Bregman distance.
A new metric MSD detects bias in datasets efficiently.
problem Detecting bias in AI systems and datasets.
method Introduced Maximum Subgroup Discrepancy (MSD) metric and a practical algorithm based on MIO.
result MSD provides a linear sample complexity for practical applications, distinguishing biases effectively.
Unified curvature for hypergraphs from Ollivier-Ricci.
problem Generalizing curvature to hypergraphs.
method Developed ORCHID framework to generalize Ollivier-Ricci curvature to hypergraphs.
result ORCHID curvatures have favorable theoretical properties and are scalable for hypergraph tasks.
We consider the problem of approximate joint triangularization of a set of noisy jointly diagonalizable real matrices. Approximate joint triangularizers are commonly used in the estimation of the joint eigenstructure of a set of matrices, with applications in signal processing, linear algebra, and tensor decomposition.…
Establishes a link between heat diffusion and manifold distances in data.
problem No theoretical link between diffusion-based manifold learning and geodesic distances.
method Formulates heat geodesic embeddings based on Riemannian geometry.
result Method outperforms state-of-the-art in preserving manifold distances and cluster structure.
Proposes a new graph kernel framework using regularized Wasserstein distances.
problem Learning optimal transport distances for graph kernels.
method Introduces Regularized Wasserstein (RW) discrepancy with two regularization terms.
result Empirically validated method outperforms state-of-the-art methods.
HD-BWDM improves clustering validation in high-dimensional data.
problem Determining the right number of clusters in high-dimensional data.
method HD-BWDM integrates random projection, PCA, trimmed clustering, and medoid-based distances.
result HD-BWDM remains stable and interpretable under high-dimensional projections and contamination.
PolyGraph Discrepancy improves graph generative model evaluation.
problem Inability of existing metrics to provide an absolute performance measure and comparability across different graph descriptors.
method Approximates Jensen-Shannon distance using binary classifiers trained to distinguish between real and generated graphs.
result PGD provides a more robust and insightful evaluation compared to MMD metrics.
Agent-based model compares different COVID-19 testing policies and their effectiveness.
problem Understanding how different testing policies reveal the true number of infected cases.
method Developed an agent-based simulation framework in Python to model various testing policies and interventions.
result Contact Tracing consistently captures more positive cases than Random Symptomatic Testing, and LBT performs similarly.
The reliable measurement of confidence in classifiers' predictions is very important for many applications and is, therefore, an important part of classifier design. Yet, although deep learning has received tremendous attention in recent years, not much progress has been made in quantifying the prediction confidence of…