We extend the Fourier cosine method to discrete probability distributions, achieving faster convergence rates.
problem Extending Fourier cosine method to discrete probability distributions.
method Spectral filters and convergence rates analysis.
result Spectral filters achieve one order faster convergence rates than previously recognized.
We analyze the probability of ruin for the {\it scaled} classical Cramér-Lundberg (CL) risk process and the corresponding diffusion approximation. The scaling, introduced by Iglehart \cite{I1969} to the actuarial literature, amounts to multiplying the Poisson rate $\la$ by n, dividing the claim severity by $\sqrtn$, …
The paper improves the probability flow ODE sampler for faster sampling of natural images.
problem Improving the convergence rate of the probability flow ODE sampler.
method Adapting the probability flow ODE sampler to exploit intrinsic low-dimensional structures in natural image data.
result Achieves a dimension-free convergence rate of O(k/T) in total variation distance, improving upon existing results. In this paper, we study stochastic non-convex optimization with non-convex random functions. Recent studies on non-convex optimization revolve around establishing second-order convergence, i.e., converging to a nearly second-order optimal stationary points. However, existing results on stochastic non-convex optimizatio…
Paper analyzes convergence of ODE samplers in Wasserstein distances.
problem Limited theoretical understanding of convergence properties of probability flow ODEs.
method Convergence analysis for general probability flow ODEs in 2-Wasserstein distance.
result First non-asymptotic convergence analysis for probability flow ODE samplers.
Study proves convergence of interest rate model approximations.
problem Investigating convergence of stochastic interest rate models.
method Developed analytical tools for true and truncated EM solutions, proving convergence in probability.
result True solution converges in probability to truncated EM solution as step size approaches zero.
SGD converges with positive probability for non-convex deep neural networks under specific conditions.
problem Convergence of SGD for non-convex deep neural networks.
method Established local convergence with positive probability under local Łojasiewicz condition and additional structural assumption.
result SGD converges with positive probability for non-convex deep neural networks under specific conditions.
The paper studies PCA of probability measures with varying sample sizes and finds optimal convergence rates.
problem PCA of multiple probability measures with varying sample sizes.
method Double asymptotic regime analysis with convergence rates n−1/2+m−α for empirical covariance and PCA risk. result Optimal convergence rates for empirical covariance and PCA risk in the dense regime are proven.
SGD converges almost surely in non-convex problems, avoiding saddle points and accelerating convergence.
problem Understanding convergence of SGD in non-convex optimization problems.
method Analysis of SGD trajectories, focusing on boundedness, convergence to strict saddle points, and rate of convergence.
result SGD converges almost surely to a minimizer in non-convex problems, avoiding strict saddle points.
The study examines how shallow neural nets converge to training samples or manifold points during diffusion.
problem Understanding when and how shallow neural nets converge to training samples or manifold points during diffusion.
method Analysis of shallow ReLU neural network denoisers trained with minimal ℓ2 norm, comparing score flow and diffusion flow. result Probability flow converges to training points, sums of training points, or manifold points, depending on the diffusion time scheduler.
RS-NSGD improves SGD convergence for heavy-tailed noise.
problem Nonconvex optimization with heavy-tailed noise.
method Integrates direction normalization into subspace updates.
result Achieves better oracle complexity than full-dimensional normalized SGD.
The paper shows how MMD metrizes weak convergence for certain kernels.
problem Characterizing MMD metrizing weak convergence for a wide class of kernels.
method Proving MMD metrizes weak convergence for specific kernels on a locally compact space.
result Corrected prior results and identified new kernels metrizing weak convergence.
Paper analyzes high probability convergence of adaptive SGD with momentum.
problem Theoretical understanding of adaptive SGD with momentum in nonconvex settings is incomplete.
method High probability analysis under weak assumptions.
result First high probability convergence proof for gradients to zero in Delayed AdaGrad with momentum.
For a sequence of nonnegative random variables, we provide simple necessary and sufficient conditions to ensure that each sequence of its forward convex combinations converges in probability to the same limit. These conditions correspond to an essentially measure-free version of the notion of uniform integrability.
Paper proves convergence of Gini index to equilibrium in Wasserstein distance.
problem Proving convergence of Gini index to equilibrium in Wasserstein distance.
method Analyzes Gini index as Lyapunov functional and proves convergence in Wasserstein distance.
result Proves convergence of Gini index to equilibrium in Wasserstein distance.
Policy gradient converges linearly with Hadamard parameterization in tabular settings.
problem Convergence of policy gradient methods under Hadamard parameterization.
method Studied convergence rate and established linear convergence after k0 iterations. result Algorithm converges linearly with rate $O(rac{1}{k})$ and faster locally after k0. This study analyzes AdaGrad's stability and convergence in non-convex optimization.
problem Lack of theoretical analysis for AdaGrad in non-convex optimization.
method Novel stopping time-based techniques from probability theory.
result Established stability and derived convergence rates for AdaGrad.
The paper analyzes convergence of Langevin dynamics with time-dependent metrics.
problem Analyzing convergence of Langevin dynamics with time-dependent metrics.
method Formulated a modified gradient flow of the Kullback-Leibler divergence, selected a time-dependent relative Fisher information functional, and developed a time-dependent Hessian matrix condition.
result Proved convergence conditions for various Langevin dynamics.
Geometric tempering improves sampling from distributions, with exponential convergence rates.
problem Sampling from probability distributions using gradient flow dynamics.
method Geometric tempering of the target distribution in Wasserstein and Fisher-Rao gradient flows.
result Exponential convergence in continuous and discrete time for geometric tempering.
Adam converges with high probability under unconstrained non-convex smooth stochastic optimizations.
problem Theoretical limitations of Adam's convergence under unconstrained non-convex smooth stochastic optimizations.
method Deep analysis of Adam's convergence rate under affine variance noise, without bounded gradient assumptions.
result Adam converges to the stationary point with a high probability rate of $\mathcal{O}\left({
m poly}(\log T)/\sqrt{T}
ight)$.
In this work we investigate to which extent one can recover class probabilities within the empirical risk minimization (ERM) paradigm. The main aim of our paper is to extend existing results and emphasize the tight relations between empirical risk minimization and class probability estimation. Based on existing literat…
New bounds for generative models under weaker assumptions.
problem Establishing convergence guarantees for generative models under weak assumptions.
method Non-asymptotic 2-Wasserstein distance bounds for probability flow ODEs under weak log-concavity and Lipschitz continuity.
result Concrete convergence rates for generative models, including non-log-concave distributions.
The CEV model is given by the stochastic differential equation Xt=X0+∫0tμXsds+∫0tσ(Xs+)pdWs, 21≤p<1. It features a non-Lipschitz diffusion coefficient and gets absorbed at zero with a positive probability. We show the weak convergence of Euler-Maruyama approximations Xtn to the proc…
New theory improves diffusion model convergence for generating data.
problem Improving convergence of diffusion models for data generation.
method Developed a non-asymptotic convergence theory for probability flow ODEs.
result Proves d/ε iterations suffice for approximating target distributions. Portfolio turnpikes state that, as the investment horizon increases, optimal portfolios for generic utilities converge to those of isoelastic utilities. This paper proves three kinds of turnpikes. In a general semimartingale setting, the abstract turnpike states that optimal final payoffs and portfolios converge under …
New algorithm FLUTE achieves uniform-PAC convergence in RL with linear approx.
problem RL with linear function approximation lacks uniform-PAC guarantees.
method FLUTE algorithm with minimax value function estimator and multi-level partition scheme.
result Uniform-PAC convergence to optimal policy with high probability.
We study the fundamental problem of learning an unknown, smooth probability function via pointwise Bernoulli tests. We provide a scalable algorithm for efficiently solving this problem with rigorous guarantees. In particular, we prove the convergence rate of our posterior update rule to the true probability function in…
In many signal detection and classification problems, we have knowledge of the distribution under each hypothesis, but not the prior probabilities. This paper is aimed at providing theory to quantify the performance of detection via estimating prior probabilities from either labeled or unlabeled training data. The erro…
MPF method improves parameter estimation in probabilistic models.
problem Difficulty in fitting probabilistic models due to intractable partition function.
method Minimum Probability Flow (MPF) method for parameter estimation.
result MPF outperforms existing techniques in convergence time and accuracy.
This paper provides an existence-and-uniqueness theorem characterizing the stochastic integral with respect to a Wiener process. The integral is represented as a mapping from the space of measurable and adapted pathwise locally integrable processes to the space of continuous adapted processes. It is characterized in te…
Paper analyzes convergence of stochastic methods under heavy-tailed noise.
problem Analyzing convergence of stochastic methods under heavy-tailed noise.
method Investigates vanilla and clipped stochastic subgradient descent methods.
result Demonstrates convergence properties under sub-Weibull and p-BCM noise assumptions.
This work addresses the convergence of SGD's final iterate without restrictive assumptions.
problem Prove optimal convergence rate of SGD's final iterate without compact domains or bounded noise.
method Unified proof for general domains, composite objectives, non-Euclidean norms, etc.
result First unified convergence rates in expectation and high probability.
A new dynamical formulation of log-PCA captures local principal modes of geodesic variations.
problem Learning principal variations of random probability measures under Wasserstein geometry.
method Introducing a new dynamical formulation of log-PCA as a variational approach.
result Deriving a general statistical convergence rate for empirical WT-PCA.
We introduce an evolutionary game with feedback between perception and reality, which we call the reality game. It is a game of chance in which the probabilities for different objective outcomes (e.g., heads or tails in a coin toss) depend on the amount wagered on those outcomes. By varying the `reality map', which rel…
Paper introduces GSPMs for robust probability metrics.
problem Lack of well-established convergence behavior for probability metrics.
method Introduces Generalized Sliced Probability Metrics (GSPMs) based on generalized Radon transform.
result GSPMs converge to global optimum under mild assumptions for generative modeling.
We associate certain probability measures on R to geodesics in the space $\H_L$ of positively curved metrics on a line bundle L, and to geodesics in the finite dimensional symmetric space of hermitian norms on H0(X,kL). We prove that the measures associated to the finite dimensional spaces converge weakly to t…
PWGF escapes saddle points in nonconvex optimization.
problem Escaping saddle points in nonconvex optimization.
method PWGF uses noisy perturbations via Gaussian process to escape saddle points.
result PWGF achieves second-order optimality for nonconvex objectives.
In this paper, we present a probability one convergence proof, under suitable conditions, of a certain class of actor-critic algorithms for finding approximate solutions to entropy-regularized MDPs using the machinery of stochastic approximation. To obtain this overall result, we prove the convergence of policy evaluat…
New metrics avoid high-dimensional analysis challenges, proving convergence without 'curse of dimensionality'.
problem High-dimensional analysis challenges in empirical measure convergence.
method Proposed a new class of probability metrics free of the curse of dimensionality.
result Convergence of empirical measures is free of the curse of dimensionality.
Study shows financial value of weak information converges in discrete vs continuous markets.
problem Analyzing financial value of weak information in discrete vs continuous markets.
method Defined minimal probability measure and financial value of weak information, then showed convergence.
result Financial value of weak information converges in discrete vs continuous markets.
A new algorithm for optimizing probability distributions converges linearly.
problem Optimizing functionals over families of probability distributions.
method Variational transport: particle-based algorithm approximating Wasserstein gradient descent.
result Variational transport converges linearly to the global minimum of the objective functional.
Aggregates probability models using Wasserstein space and variational approach.
problem Model aggregation in the Wasserstein space of distributions.
method Data-driven calibration framework based on Γ-convergence. result Empirical minimizers converge to the minimizers of the actual problem.
A new optimization method for probability simplex problems.
problem Optimizing convex problems over the probability simplex.
method Cauchy-Simplex iteration scheme, mapping to sphere, gradient descent, and back-mapping.
result Convergence results and faster convergence in high dimensions.
We consider the smoothing probabilities of hidden Markov model (HMM). We show that under fairly general conditions for HMM, the exponential forgetting still holds, and the smoothing probabilities can be well approximated with the ones of double sided HMM. This makes it possible to use ergodic theorems. As an applicatio…
SGD methods fail to converge to global minimizers in deep neural networks with ReLU activation.
problem Failure of SGD methods to converge to global minimizers in deep neural networks.
method Stochastic Gradient Descent (SGD) and its variants like Adam, RMSProp, etc.
result SGD methods fail to converge to global minimizers with high probability in deep neural networks with ReLU activation.
The paper improves OT map estimation rates without strict assumptions.
problem Estimating optimal transport maps under practical conditions.
method Developed new convergence rates and scalable algorithms.
result Improved convergence rates for OT map estimation without restrictive assumptions.
The probabilistic bisection algorithm (PBA) solves a class of stochastic root-finding problems in one dimension by successively updating a prior belief on the location of the root based on noisy responses to queries at chosen points. The responses indicate the direction of the root from the queried point, and are incor…
We derive the fast convergence rates of a deep neural network (DNN) classifier with the rectified linear unit (ReLU) activation function learned using the hinge loss. We consider three cases for a true model: (1) a smooth decision boundary, (2) smooth conditional class probability, and (3) the margin condition (i.e., t…