MPF method improves parameter estimation in probabilistic models.
problem Difficulty in fitting probabilistic models due to intractable partition function.
method Minimum Probability Flow (MPF) method for parameter estimation.
result MPF outperforms existing techniques in convergence time and accuracy.
The study examines how shallow neural nets converge to training samples or manifold points during diffusion.
problem Understanding when and how shallow neural nets converge to training samples or manifold points during diffusion.
method Analysis of shallow ReLU neural network denoisers trained with minimal ℓ2 norm, comparing score flow and diffusion flow. result Probability flow converges to training points, sums of training points, or manifold points, depending on the diffusion time scheduler.
We study the misclassification error for community detection in general heterogeneous stochastic block models (SBM) with noisy or partial label information. We establish a connection between the misclassification rate and the notion of minimum energy on the local neighborhood of the SBM. We develop an optimally weighte…
Paper bounds PAC RL sample complexity in deterministic MDPs.
problem Identify ε-optimal policy with high probability.
method Proposes nearly matching upper and lower bounds on sample complexity, introduces deterministic return gap, uses graph-theoretical concepts and maximum-coverage exploration.
result First nearly matching upper and lower bounds on sample complexity for PAC RL in deterministic MDPs.
New method improves MMD estimation without convexity assumptions.
problem Lack of theoretical guarantees for MMD estimation algorithms.
method Preconditioned gradient descent (PGD) scheme for MMD optimization.
result PGD scheme converges globally under specific conditions.
When maximum likelihood estimation is infeasible, one often turns to score matching, contrastive divergence, or minimum probability flow to obtain tractable parameter estimates. We provide a unifying perspective of these techniques as minimum Stein discrepancy estimators, and use this lens to design new diffusion kerne…
The quest for biologically plausible deep learning is driven, not just by the desire to explain experimentally-observed properties of biological neural networks, but also by the hope of discovering more efficient methods for training artificial networks. In this paper, we propose a new algorithm named Variational Proba…
The study tightens bounds on binomial probabilities and minimums using KL-divergence.
problem Tightening bounds on binomial probabilities and minimums of i.i.d. Binomials.
method Applied Sanov's theorem to derive upper and lower bounds on binomial tail probabilities and minimums, expressed in terms of KL-divergence.
result High probability upper and lower bounds on the minimum of i.i.d. Binomial random variables, finite sample, asymptotically tight.
Gradient flow in parameters equals linear interpolation in outputs.
problem Understanding and optimizing training algorithms in deep learning.
method Proving equivalence between gradient flow in parameter space and linear interpolation in output space, and deriving formulas for global minima.
result Gradient flow in parameters can be transformed into linear interpolation in outputs, leading to global minima.
Efficient adjustment sets found for cost-minimized causal estimations.
problem Estimating interventional means with minimum cost in causal graphical models.
method Defined cost-adjustment sets, constructed flow networks, and used maximum flow algorithms.
result Minimum cost optimal adjustment sets exist and can be found efficiently.
Study uses Bayes Hilbert framework to recover probability measure flows from sensors.
problem Recovering probability measure flows from moving sensors in a Hilbert space.
method Bayes Hilbert framework, minimum-energy transport, linearization, variational theory.
result Localized sensors can recover reduced path directions but not full state space.
Gradient flow in phase retrieval escapes spurious minima with high probability.
problem Understanding gradient-based optimization in high-dimensional non-convex functions.
method Analytical and numerical study of gradient dynamics in phase retrieval.
result Gradient flow avoids spurious minima by drifting along unstable directions.
Fitting probabilistic models to data is often difficult, due to the general intractability of the partition function and its derivatives. Here we propose a new parameter estimation technique that does not require computing an intractable normalization factor or sampling from the equilibrium distribution of the model. T…
Proposes a thermodynamic work minimization framework for guiding generative models.
problem Guiding generative models in sparse-data regimes with limited target samples or constraints.
method Regularization framework inspired by thermodynamic work, introducing Path Guidance and Observable Guidance.
result Improves sample efficiency and reduces bias in molecular simulations.
This study explains gradient flow dynamics in neural networks for small initialisation.
problem Understanding the training dynamics of neural networks for small initialisation.
method Analysis of gradient flow dynamics for one-hidden layer ReLU networks with orthogonal inputs.
result Gradient flow converges to zero loss and characterizes implicit bias towards minimum variation norm.
Associating distinct groups of objects (clusters) with contiguous regions of high probability density (high-density clusters), is central to many statistical and machine learning approaches to the classification of unlabelled data. We propose a novel hyperplane classifier for clustering and semi-supervised classificati…
Using transfer entropy, we observed the strength and direction of information flow between stock indices. We uncovered that the biggest source of information flow is America. In contrast, the Asia/Pacific region the biggest is receives the most information. According to the minimum spanning tree, the GSPC is located at…
New insights into matrix factorization show strict saddles have bounded eigenvalues.
problem Understanding the nature of critical points in matrix factorization.
method Analyzing orbits of critical points under the general linear group and identifying canonical points.
result Minimum eigenvalue of strict saddles is not uniformly bounded below zero.
We present an efficient algorithm for maximum likelihood estimation (MLE) of exponential family models, with a general parametrization of the energy function that includes neural networks. We exploit the primal-dual view of the MLE with a kinetics augmented model to obtain an estimate associated with an adversarial dua…
Study rigidity of Hamiltonians near a minimum in symplectic and magnetic settings.
problem Rigidity of Hamiltonians near a minimum in symplectic and magnetic settings.
method Analyzing Hamiltonian systems near a compact symplectic Morse-Bott minimum, focusing on Zoll flows and magnetic forms.
result A constant curvature quantity characterizes complex space forms among Kähler manifolds.
Wide neural networks with asymmetrical node scaling converge globally and learn features.
problem Global convergence and feature learning in over-parameterised shallow networks.
method Gradient-based optimisation of wide, shallow neural networks with asymmetrical node scaling.
result Gradient flow and gradient descent converge to a global minimum and learn features, unlike in the NTK parameterisation.
In this paper, we mainly study the mean curvature flow in Kähler surfaces with positive holomorphic sectional curvatures. We prove that if the ratio of the maximum and the minimum of the holomorphic sectional curvatures is less than 2, then there exists a positive constant δ depending on the ratio such that $\cosα\ge…
Identifies most probable flows for Kunita SDEs in fluid dynamics.
problem Modeling stochastic processes with Eulerian noise and deterministic drifts.
method Equipping the domain with a Riemannian metric from the noise, solving the resulting PDEs.
result Most probable flows differ from deterministic flows, especially under noise.
The paper describes fitting submanifolds to data using Sussmann's orbit theorem.
problem Fitting an immersed submanifold to random samples.
method Uses Sussmann's orbit theorem to ensure submanifold fitting. Reconstruction involves encoding times and decoding via flows of vector fields.
result A high-probability bound on excess risk for the reconstruction error.
Deep linear networks can closely approximate interpolants without improving risk.
problem Understanding the risk bounds of deep linear networks compared to minimum ℓ2-norm solutions. method Bounding excess risk of interpolating deep linear networks trained using gradient flow.
result Deep linear networks can closely approximate or match minimum ℓ2-norm solutions in terms of risk. Gradient flow converges to a minimal convex structure.
problem Finding the minimal convex structure in hyperbolic manifolds.
method Weil-Petersson gradient vector field of renormalized volume.
result The flow converges to the structure with minimum convex core volume.
Gradient flows on distributions of distributions for machine learning tasks.
problem Designing gradient flows for datasets of probability distributions.
method Representing classes as conditional distributions, modeling datasets as mixture distributions, using Wasserstein over Wasserstein (WoW) distance and gradients.
result Demonstrated gradient flows for dataset transfer and distillation tasks.
The paper improves the probability flow ODE sampler for faster sampling of natural images.
problem Improving the convergence rate of the probability flow ODE sampler.
method Adapting the probability flow ODE sampler to exploit intrinsic low-dimensional structures in natural image data.
result Achieves a dimension-free convergence rate of O(k/T) in total variation distance, improving upon existing results. New probability path model improves flow matching forecasting performance.
problem Impact of probability path model selection on flow matching forecasting performance.
method Proposed a novel probability path model designed to improve forecasting performance.
result Our model achieves faster convergence during training and improved predictive performance compared to existing models.
Paper analyzes convergence of ODE samplers in Wasserstein distances.
problem Limited theoretical understanding of convergence properties of probability flow ODEs.
method Convergence analysis for general probability flow ODEs in 2-Wasserstein distance.
result First non-asymptotic convergence analysis for probability flow ODE samplers.
In this note, we explicitly solve the problem of maximizing utility of consumption (until the minimum of bankruptcy and the time of death) with a constraint on the probability of lifetime ruin, which can be interpreted as a risk measure on the whole path of the wealth process.
Flow Matching enables robust training of CNFs with various probability paths.
problem Training Continuous Normalizing Flows (CNFs) at large scales.
method Flow Matching (FM) is a simulation-free approach for training CNFs by regressing vector fields of conditional probability paths.
result Flow Matching with diffusion paths yields more robust and stable training compared to diffusion-based methods.
NOFIS uses normalizing flows to estimate rare event probabilities more efficiently.
problem Accurate estimation of rare event probabilities using conventional methods is inefficient and resource-intensive.
method NOFIS learns a sequence of proposal distributions by minimizing KL divergence losses and estimates rare event probability using importance sampling.
result NOFIS outperforms baseline approaches in estimating rare event probabilities across 10 distinct test cases.
Gradient descent with geometrically adapted metrics drives L2 cost to global minimum at uniform rate.
problem Minimizing L2 cost in deep learning networks. method Adapting gradient descent to output layer metric in deep learning.
result Uniform exponential convergence to global minimum in L2 cost. Bi-Lipschitz flows approximate a wide range of distributions.
problem Characterizing the expressivity of bi-Lipschitz normalizing flows.
method Linking score regularity to transport map bi-Lipschitzness via probability flow ODE.
result Gaussian pullbacks induced by bi-Lipschitz variance-preserving transport maps are L1-dense among all probability densities. The article analyzes the stability of a curve shortening flow for planar networks.
problem Stability analysis of anisotropic curve shortening flow for planar networks.
method Used Lojasiewicz-Simon gradient inequality to derive stability results.
result For initial data close to an energy minimizer, the flow exists globally and converges to a different energy minimum.
While the channel capacity reflects a theoretical upper bound on the achievable information transmission rate in the limit of infinitely many bits, it does not characterise the information transfer of a given encoding routine with finitely many bits. In this note, we characterise the quality of a code (i. e. a given en…
A new generative model relaxes the bijectivity requirement for invertible flows.
problem Challenges of invertible flow-based models in scaling to large datasets.
method Proposes a generative model based on relaxed injective probability flows.
result Improves sample quality over VAEs and AEs.
This paper studies gradient flows for sampling using various metrics and their affine invariance.
problem Sampling from probability distributions with unknown normalizations.
method Gradient flows in the space of probability measures, focusing on Kullback-Leibler divergence and affine invariance of metrics.
result Gradient flows of Kullback-Leibler divergence do not depend on the normalization constant, and affine invariance is achieved for certain metrics.
In this paper, we define a certain "proportional volume property" for an unit vector field on a spherical domain in S3. We prove that the volume of these vector fields has an absolute minimum and this value is equal to the volume of the Hopf vector field. Some examples of such vector fields are given. We also study the…
We propose a new topic modeling procedure that takes advantage of the fact that the Latent Dirichlet Allocation (LDA) log likelihood function is asymptotically equivalent to the logarithm of the volume of the topic simplex. This allows topic modeling to be reformulated as finding the probability simplex that minimizes …
The development of algorithms for unsupervised pattern recognition by nonlinear clustering is a notable problem in data science. Markov clustering (MCL) is a renowned algorithm that simulates stochastic flows on a network of sample similarities to detect the structural organization of clusters in the data, but it has n…
A new method for sampling from complex distributions using Langevin samplers.
problem Sampling from unnormalized Boltzmann densities.
method Probability flow ODE derived from linear stochastic interpolants, employing Langevin samplers.
result Efficient simulation of the flow with non-asymptotic convergence rate.
Flow-based models use ODEs to generate complex data distributions.
problem Generating high-dimensional data with complex probability distributions.
method Flow-based models use invertible mappings governed by ODEs to capture these distributions.
result Flow-based models provide exact likelihood estimation and efficient sampling.
New error bounds for flow matching methods using deterministic sampling.
problem Improving the accuracy of flow matching methods for generating probability distributions.
method Derived error bounds for flow matching methods under deterministic sampling conditions.
result Presented error bounds for flow matching methods using L2 loss and regularity conditions. In this paper, we use the distance comparison principle, first been developed by G. Huisken, to study the spatial curve shortening flow. We have got the result that if the initial curve is the helix, then the local minimum of the ratio of the extrinsic and intrinsic distance is non-decreasing. And we have proved a Gray…
We study the risk of minimum-norm interpolants of data in Reproducing Kernel Hilbert Spaces. Our upper bounds on the risk are of a multiple-descent shape for the various scalings of d=nα, α∈(0,1), for the input dimension d and sample size n. Empirical evidence supports our finding that minimum-norm interpo…
Gradient descent converges to a global minimum in nonlinear ReLU implicit networks with linear width.
problem Understanding convergence of gradient methods in nonlinear, infinitely deep ReLU networks.
method Introduced a scaling constant to ensure well-posedness of the equilibrium equation, proving convergence to a global minimum for linear width networks.
result Gradient descent converges to a global minimum at a linear rate for nonlinear ReLU implicit networks with linear width.