CNFs learn on manifolds using PPD, improving likelihood and sample quality.
problem Training CNFs on manifolds efficiently and accurately.
method Minimizing PPD, a novel divergence, to train CNFs on manifolds.
result CNFs trained with PPD achieve state-of-the-art results on manifold benchmarks.
New analysis of annealing paths in sampling and estimation.
problem Sampling from complex distributions and estimating normalization constants.
method Extending known results on Bregman divergence to quasi-arithmetic means under monotonic embedding.
result Analogous result for quasi-arithmetic means, highlighting the interplay between means, parametric families, and divergence functionals.
Improved KL divergence estimators for normalizing flows lead to faster convergence and better approximations.
problem Estimating KL divergences for normalizing flows efficiently and accurately.
method Path-gradient estimators for reverse and forward KL divergences.
result Path-gradient estimators lead to faster convergence and better approximation results.
Sparse RSP routing improves graph exploration and classification.
problem Optimal randomized routing and distance measures on weighted graphs.
method Tsallis divergence regularization for sparse RSP.
result Sparse random walk converges to least-cost graph as temperature decreases.
We consider the problem of function approximation by two-layer neural nets with random weights that are "nearly Gaussian" in the sense of Kullback-Leibler divergence. Our setting is the mean-field limit, where the finite population of neurons in the hidden layer is replaced by a continuous ensemble. We show that the pr…
This paper bridges variational inference and Wasserstein gradient flows.
problem Combining variational inference and Wasserstein gradient flows for more efficient approximations.
method Recasting Bures-Wasserstein gradient flow as a Euclidean gradient flow and using path-derivative gradient estimator.
result A new gradient estimator for f-divergences that can be implemented using machine learning libraries. The paper develops divergences for Gaussian processes and RKHS settings.
problem Estimating divergences in infinite-dimensional spaces.
method Formulations of Alpha Log-Det divergences, continuity in norm, laws of large numbers, consistent estimation from finite samples.
result Infinite-dimensional divergences can be estimated from finite-dimensional versions with dimension-independent sample complexities.
New findings show score matching's accuracy doesn't ensure numerical stability in diffusion sampling.
problem Numerical stability issues in diffusion sampling despite small forward-marginal error.
method Constructing a smooth score field with arbitrarily small forward-marginal L2 error, showing nonexplosive behavior and moments of every order. result Euler--Maruyama discretizations can converge in probability even when moments diverge, demonstrating failure of weak convergence.
Recognizing subtle historical patterns is central to modeling and forecasting problems in time series analysis. Here we introduce and develop a new approach to quantify deviations in the underlying hidden generators of observed data streams, resulting in a new efficiently computable universal metric for time series. Th…
The study explores geodesics and KL-divergence on Hölder equilibrium probabilities.
problem Finding the probability that minimizes KL-divergence from a fixed probability in a convex set of probabilities.
method Analyzes geodesics paths on the manifold of Hölder equilibrium probabilities and uses KL-divergence as a metric.
result Explicit equations for the solution of the minimization problem are derived.
The paper explores statistical and topological properties of sliced probability divergences.
problem Understanding the topological, statistical, and computational consequences of slicing divergences.
method Deriving theoretical properties of sliced probability divergences, including metric axioms preservation and weak continuity.
result Sliced divergences share similar topological properties and have stable sample complexity.
Introduces q-paths for generalizing geometric annealing paths in machine learning.
problem Limited applicability of existing path methods in machine learning.
method Develops a family of paths derived from a generalized mean, including geometric and arithmetic mixtures.
result Empirical gains in Bayesian inference and generative model evaluation.
Develops a new divergence framework that combines f-divergences and IPMs.
problem Comparing distributions that are not absolutely continuous.
method Introduces (f,Γ)-divergences as a two-stage mass-redistribution/mass-transport process. result Improves estimation, learning, and uncertainty quantification in GANs for heavy-tailed distributions.
Study of most probable paths for anisotropic Brownian motions on manifolds.
problem Characterizing paths of Brownian motions with anisotropic diffusion on manifolds.
method Using stochastic development and fiber bundle of linear frames, the study provides a comprehensive characterization of most probable paths.
result Explicit equations and integration methods for most probable paths on different geometries, including constant curvature surfaces.
The path probability of a particle undergoing stochastic motion is studied by the use of functional technique, and the general formula is derived for the path probability distribution functional. The probability of finding paths inside a tube/band, the center of which is stipulated by a given path, is analytically eval…
New Wasserstein divergence improves generative model robustness and structure preservation.
problem Improving generative model robustness and structure preservation.
method Introduces a novel Wasserstein-1 path-space divergence and a WUP theorem.
result Derives robustness and generalization bounds for flow-based models.
New probability path model improves flow matching forecasting performance.
problem Impact of probability path model selection on flow matching forecasting performance.
method Proposed a novel probability path model designed to improve forecasting performance.
result Our model achieves faster convergence during training and improved predictive performance compared to existing models.
Understanding and measuring model risk is important to financial practitioners. However, there lacks a non-parametric approach to model risk quantification in a dynamic setting and with path-dependent losses. We propose a complete theory generalizing the relative-entropic approach by Glasserman and Xu to the dynamic ca…
f-divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler divergence, chi-squared divergence, squared Hellinger distance, total variation distance e…
Analyzed a generative model framework through Wasserstein Gradient Flow.
problem Generative modeling challenges.
method Wasserstein Gradient Flow (WGF) interpretation of Drifting Models (GMD).
result Different algorithms correspond to specific limiting points of WGFs on various divergences.
This study examines biases in flow matching samplers using finite-sample estimation.
problem Biases in flow matching samplers when using finite-sample surrogates.
method Finite-sample plug-in estimation and hierarchy of empirical FM models.
result Exact empirical minimizer and smoothed plug-in regime identified for affine conditional flows.
Optimizes diffusion processes for target distributions.
problem Efficiently generating target distributions from point masses.
method Stochastic interpolant framework with conditional expectation drift.
result Optimal diffusion coefficient minimizes path-space KL divergence.
New AI method generates SDE paths without explicit coefficients.
problem Simulating unknown Markovian SDEs with limited data.
method Uses conditional diffusion models on sample paths.
result Consistently outperforms alternative methods in KL divergence.
Develops methods to find most probable paths on complex manifolds.
problem Identifying optimal paths for manifold-valued processes, especially those with non-trivial structures.
method Constructs a general approach to defining and identifying most probable paths by measuring the Onsager-Machlup function on the anti-development of such processes.
result Derives explicit equations for development most probable paths that encompass various manifold-valued processes.
Paper introduces symmetric divergence link models for probability distributions.
problem Symmetric divergence measures for probability distributions.
method Two general classes of link models: one for survival functions and another for cumulative probability distribution functions.
result Advantages of symmetric divergence measures over asymmetric measures for model averaging and feature assessment.
Flow Matching enables robust training of CNFs with various probability paths.
problem Training Continuous Normalizing Flows (CNFs) at large scales.
method Flow Matching (FM) is a simulation-free approach for training CNFs by regressing vector fields of conditional probability paths.
result Flow Matching with diffusion paths yields more robust and stable training compared to diffusion-based methods.
Unified view of KL-divergence and IPMs via DRE, with new DRM metrics.
problem Unified understanding of KL-divergence and IPMs.
method Unified representation via maximum likelihood density-ratio estimation (DRE).
result Unified form of IPMs and novel DRM metrics.
Sharp bounds for high-probability estimation of discrete distributions.
problem Estimating discrete distributions with high probability under χ2-divergence. method Sharp upper and lower bounds for the classical Laplace estimator, and characterization of minimax high-probability risk for any estimator.
result Sharp bounds for high-probability estimation of discrete distributions can be achieved through a simple smoothing strategy.
The study tightens bounds on binomial probabilities and minimums using KL-divergence.
problem Tightening bounds on binomial probabilities and minimums of i.i.d. Binomials.
method Applied Sanov's theorem to derive upper and lower bounds on binomial tail probabilities and minimums, expressed in terms of KL-divergence.
result High probability upper and lower bounds on the minimum of i.i.d. Binomial random variables, finite sample, asymptotically tight.
This work develops a generic framework, called the bag-of-paths (BoP), for link and network data analysis. The central idea is to assign a probability distribution on the set of all paths in a network. More precisely, a Gibbs-Boltzmann distribution is defined over a bag of paths in a network, that is, on a representati…
CR-AIS improves AIS efficiency by constant rate annealing.
problem Efficiently sample from intractable distributions.
method Constant rate annealing schedule for AIS.
result CR-AIS outperforms existing Adaptive AIS methods.
New method uses neural networks to solve complex PDEs from optimal control theory.
problem Solving high-dimensional Hamilton-Jacobi-Bellman PDEs.
method Iterative diffusion optimization techniques, focusing on path measures and divergences.
result Favourable properties of log-variance divergence for Monte Carlo estimators.
We show that both Teichmuller space (with the Teichmuller metric) and the mapping class group (with a word metric) have geodesic divergence that is intermediate between the linear rate of flat spaces and the exponential rate of hyperbolic spaces. For every two geodesic rays in Teichmuller space, we find that their dive…
Proposes a new divergence measure for probability distributions.
problem Challenges in estimating divergences from empirical samples.
method Embeds data into RKHS, computes Jensen-Shannon divergence between covariance operators.
result Establishes RJSD as a lower bound on Jensen-Shannon divergence, enabling variational estimation.
We present an explicit realization of abelian extensions of infinite dimensional Lie groups using abelian extensions of path groups, by generalizing Mickelsson's approach to loop groups and the approach of Losev-Moore-Nekrasov-Shatashvili to current groups. We apply our method to coupled cocycles on current Lie algebra…
Temporal aggregation reveals latent default correlation from monthly data.
problem Understanding effective default correlation from monthly default data.
method Temporal coarse-graining of latent default-probability paths.
result Temporal coarse-graining improves identifiability and reduces over-allocation of long-horizon fluctuations.
The paper explores geometry of probability measures and barycenter maps.
problem Understanding the space of probability measures and their barycenter.
method Information geometry, Fisher metric, dualistic structures, divergences, geodesics.
result Recent developments in the geometry of probability measures and barycenter.
Construction of ambiguity set in robust optimization relies on the choice of divergences between probability distributions. In distribution learning, choosing appropriate probability distributions based on observed data is critical for approximating the true distribution. To improve the performance of machine learning …
Temporal coarse-graining of latent default paths explains effective correlation in corporate defaults.
problem Understanding effective default correlation in corporate defaults.
method Temporal coarse-graining of latent default-probability paths, applied to corporate default-count data.
result Temporal coarse-graining provides a scale-consistent baseline that improves identifiability and reduces over-allocation of long-horizon fluctuations.
Theoretical proof shows COMs are a type of contrastive divergence model with improved sampling.
problem Improving sampling quality in offline model-based optimization.
method Showed COMs are contrastive divergence models, proposed Langevin MCMC sampler, and decoupled model.
result Improved sampling quality achieved by decoupling model and using Langevin MCMC.
Paper examines stability of Bayesian posterior measures using integral probability metrics.
problem Stability of Bayesian inference in large-scale inverse problems.
method New families of integral probability metrics for likelihood and prior perturbations.
result Constructs new stability results for Bayesian posterior measures.
We propose a general strategy to derive null-homotopy operators for differential complexes based on the Bernstein-Gelfand-Gelfand (BGG) construction and properties of the de Rham complex. Focusing on the elasticity complex, we derive path integral operators P for elasticity satisfying $\mathscr{D}\mathscr{P…
A Bayesian framework models dynamic probability predictions over time.
problem Dynamic probability predictions over time in various settings.
method Gaussian latent information martingale (GLIM) framework.
result GLIM outperforms baseline methods in predicting future uncertainties.
In this report, we derive a non-negative series expansion for the Jensen-Shannon divergence (JSD) between two probability distributions. This series expansion is shown to be useful for numerical calculations of the JSD, when the probability distributions are nearly equal, and for which, consequently, small numerical er…
Let x denote a diffusion process defined on a closed compact manifold. In an earlier article, the author introduced a new approach to constructing admissible vector fields on the associated space of paths, under the assumption of ellipticity of x. In this article, this method is extended to yield similar results fo…
Develops a machine learning framework for computing most probable paths in stochastic systems.
problem Computing the most probable paths in stochastic dynamical systems.
method Reformulates the boundary value problem of Hamiltonian systems and uses a neural network to solve the Euler-Lagrange equation for the Onsager-Machlup action functional.
result Demonstrates the efficacy and accuracy of the machine learning approach in computing most probable paths for stochastic systems with various types of noise.
The paper defines conditions for Gaussian process sample path regularity.
problem Lack of understanding of Gaussian process sample path regularity.
method Analyzes covariance kernels to determine sample path regularity.
result Necessary and sufficient conditions for Hölder regularity are provided.
BWFlow improves graph generation by smoothly interpolating graph components.
problem Disjoint modeling of graph nodes and edges leads to irregular and non-smooth probability paths.
method Modeling graphs as MRFs and using optimal transport displacement for a smooth probability path.
result BWFlow achieves better training convergence and efficient sampling in graph generation.