New bounds for heavy-tailed SDEs without info-theory terms.
problem Understanding generalization of heavy-tailed stochastic optimization.
method Fractional Fokker-Planck equation to estimate entropy flows.
result High-probability bounds with better dimension dependence.
New RDP guarantees for heavy-tailed SDEs and SGD.
problem Characterizing differential privacy for heavy-tailed noise in learning algorithms.
method Rényi flow computations and fractional Poincaré inequalities.
result First RDP guarantees for heavy-tailed SDEs with weaker dependence on dimension.
Work on SGDm under heavy-tailed noise, revealing its generalization properties.
problem Understanding generalization of SGDm under heavy-tailed noise.
method Analysis of continuous-time limit (SDE) and discrete-time SGDm, establishing generalization bounds.
result SGDm can have worse generalization in the presence of heavy-tailed noise for quadratic loss functions.
DLPM replaces Gaussian noise with α-stable noise in DDPM, improving data distribution coverage and robustness.
problem Handling mode collapse and class imbalance in datasets with heavy-tailed noise.
method Extending DDPM to use α-stable noise, simplifying the process with elementary proof techniques.
result DLPM yields better coverage of data distribution tails, improved robustness to unbalanced datasets, and faster computation times.
Self-regulating annealing improves sampling from heavy-tailed datasets.
problem Sampling from heavy-tailed distributions using diffusion models.
method Proposed an SDE-based sampler with a state-dependent diffusion coefficient.
result State dependence induces a self-regulating annealing mechanism.
New method for Bayesian inference of Lévy-driven SDEs with jumps.
problem Bayesian inference for Lévy-driven SDEs is challenging due to discontinuities and heavy tails.
method Neural exponential tilting framework for variational inference.
result Accurately captures jump dynamics and reliable posterior inference in heavy-tailed regimes.
New method handles complex systems with discontinuous, heavy-tailed noise.
problem Handling discontinuous, heavy-tailed Lévy noise in stochastic systems.
method Developed nonlocal Kramers-Moyal formulas for SDEs with multiplicative Lévy noise.
result Validated framework for discovering interpretable SDE models from data.
New bounds link SGD's generalization to heavy tails without topological assumptions.
problem Linking SGD's generalization error to heavy tails without additional assumptions.
method Developed Wasserstein stability bounds for heavy-tailed SDEs and their discretizations, converting to generalization bounds.
result Generalization bounds for a broader class of objective functions, including non-convex functions, without topological assumptions.
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the \emph{classical} central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochast…
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the classical central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochastic diff…
Proves generalization bounds for SGD using Feller processes and Hausdorff dimension.
problem Characterizing generalization properties of SGD in deep learning.
method Proves generalization bounds for SGD under Feller process approximation, linking generalization error to the Hausdorff dimension of trajectories.
result Generalization error controlled by the Hausdorff dimension of trajectories, which is linked to the tail behavior of the driving process.
Gradient descent with chaotic perturbations improves generalization.
problem Improving generalization of gradient descent.
method Introducing chaotic perturbations to gradient descent to achieve improved generalization.
result Gradient descent with chaotic perturbations converges to a heavy-tailed SDE, leading to improved generalization.
Unified framework for distributed compressed SGD under (L0,L1)-smoothness.
problem Understanding the joint effect of batch noise, adaptivity, and compression in distributed stochastic optimization.
method Developed a unified theoretical framework using SDEs that incorporate curvature-dependent terms.
result Normalizing updates in DCSGD stabilizes convergence, with normalization degree determined by noise structure and landscape regularity.
This research explains why SGD generalizes better than ADAM in deep learning.
problem Understanding the generalization gap between SGD and ADAM in deep learning.
method Analyzing local convergence behaviors through Levy-driven stochastic differential equations (SDEs).
result SGD is more locally unstable and better escapes from sharp minima to flatter ones, leading to better generalization.
Stochastic gradient descent (SGD) has been widely used in machine learning due to its computational efficiency and favorable generalization properties. Recently, it has been empirically demonstrated that the gradient noise in several deep learning settings admits a non-Gaussian, heavy-tailed behavior. This suggests tha…
The paper tackles drift identification in Lévy α-stable stochastic systems, proposing a Fourier space approach.
problem Estimating the drift field of a stochastic differential equation driven by Lévy α-stable noise.
method Fourier space approach, parameterizing the drift field using Fourier series, minimizing a loss function with gradients computed via the adjoint method.
result The method is capable of learning drift fields in qualitative and/or quantitative agreement with ground truth fields.
SDE Matching eliminates simulation for training Latent SDEs, achieving similar performance.
problem Training Latent SDEs with adjoint sensitivity methods is computationally expensive and limited.
method SDE Matching, inspired by Score- and Flow Matching, eliminates simulation for training Latent SDEs.
result SDE Matching achieves performance comparable to adjoint sensitivity methods while reducing computational complexity.
New SDEs use G-Brownian motion, extending mean-field models.
problem Extending mean-field models to new types of stochastic processes.
method Introduced G-SDEs with coefficients dependent on current state and solution as random variable. result Validated new SDE framework for complex stochastic systems.
Study on the smoothness of solutions to a specific type of stochastic differential equation.
problem Regularity of solutions to mean-field G-SDEs. method Analysis of first and second order Fréchet differentiability in the random initial condition.
result Established the Fréchet differentiability of the solution and specified the corresponding equations.
New SDE model for continuous-time reinforcement learning.
problem Modeling exploration in continuous-time reinforcement learning.
method Introduced grid-sampling SDE as a proxy model.
result Wellposedness of the SDE in the presence of jumps.
This paper uses SDEs to analyze GANs training and long-run behavior.
problem Understanding the training process and long-run behavior of GANs.
method Established SDE approximations for GANs training and analyzed long-run behavior via invariant measures.
result The long-run behavior of GANs training can be studied via the invariant measures of its SDE approximations.
The paper identifies generators of linear SDEs with noise types.
problem Identifying the generator of linear SDEs from their solution distribution.
method Deriving sufficient and necessary conditions for additive noise, and sufficient conditions for multiplicative noise.
result Generic conditions for identifying the generator of linear SDEs with both types of noise.
Characterizes term structure models driven by Lévy processes.
problem Modeling non-negative short rates with Lévy processes.
method Analyzes affine term structure models driven by independent Lévy martingales.
result All possible solutions of the models can be obtained using stable processes.
Simulation-free VI closes the approximation gap in latent SDEs
problem Recovering dynamical systems from noisy observations
method Helmholtz-SDE
result Recovers dynamics more faithfully than prior methods
Sig-SDE model integrates signatures with SDEs for financial data.
problem Calibrating models to exotic financial products with non-linear dependencies.
method Integrating signatures from stochastic analysis with neural SDEs.
result Sig-SDE provides theoretical guarantees for convergence.
Proposes methods to include distributional information in MV-SDEs for better modeling of interacting particle systems.
problem Modeling the behavior of an infinite number of interacting particles with distributional information.
method Semi-parametric methods and estimators for MV-SDEs.
result Explicitly including distributional dependence improves performance in modeling temporal data with interaction.
LatentFlow simplifies conditioning of stochastic processes without training.
problem Intractable conditional laws for complex stochastic models.
method Writing stochastic process as latent innovation, reducing conditioning to latent-space inference.
result Exact conditional sampling across various model classes.
New method learns SDEs without integrators, speeding up computation.
problem Computational expense in learning SDEs using neural networks.
method Importance-sampling estimator for SDEs, leveraging parallelism.
result Lower-variance gradient estimates and massive computation time reductions.
Stochastic normalizing flows use SDEs for efficient training and sampling.
problem Efficient maximum likelihood estimation and variational inference.
method Continuous normalizing flows extended with stochastic differential equations (SDEs) and rough path theory.
result Stochastic normalizing flows enable efficient training and sampling from complex distributions.
We explain how Itô Stochastic Differential Equations (SDEs) on manifolds may be defined using 2-jets of smooth functions. We show how this relationship can be interpreted in terms of a convergent numerical scheme. We show how jets can be used to derive graphical representations of Itô SDEs. We show how jets can be used…
New method speeds up SDE inference by matching moments to FPK equation.
problem Efficiency of sampling schemes in high-dimensional SDEs.
method Direct approximation of Fokker-Planck-Kolmogorov equation by matching moments.
result Fast, scalable inference in high-dimensional latent spaces.
Proposes neural SDEs with change points for better time series modeling.
problem Restrictions in modeling time series with distributional shift.
method Generative adversarial networks (GANs) for SDEs and change point detection.
result Jointly learns change points and SDE model parameters.
New geometric SDEs and discretizations on Riemannian manifolds with error bounds.
problem Modeling diffusion processes on Riemannian manifolds with geometric SDEs.
method Introduced a new construction of geometric SDEs and provided non-asymptotic error bounds.
result First non-asymptotic error bound for geometric Euler-Murayama discretization.
This paper bridges the gap between ODE and SDE in diffusion models using Fokker-Planck equations.
problem Empirical evidence shows that ODE-based samples from score-based diffusion models are inferior to SDE-based samples.
method The paper rigorously describes dynamics and approximations in training score-based diffusion models, linking them to Fokker-Planck equations.
result Adding a regularisation term based on the Fokker-Planck residual can close the gap between ODE- and SDE-induced distributions.
New method identifies SDE drift and diffusion from temporal data.
problem Learning SDE parameters from temporal data, especially in noisy or incomplete data.
method Entropy-regularized optimal transport, APPEX algorithm.
result Can almost always recover drift and diffusion from temporal marginals.
New algorithm optimizes nonlinear SDEs online with convergence guarantees.
problem Optimizing nonlinear stochastic differential equations (SDEs) is computationally challenging.
method Forward propagation algorithm that solves an SDE derived using forward differentiation.
result Convergence theorem for nonlinear dissipative SDEs with bounds on stochastic fluctuations.
We developed efficient methods to compute gradients for Neural SDEs, improving training speed and accuracy.
problem Training Neural SDEs requires accurate and efficient computation of gradients, which is challenging due to the complexity of SDEs.
method We introduced a reversible Heun method for solving backwards-in-time SDEs and a Brownian Interval for sampling and reconstructing Brownian motion.
result Our methods significantly improve training speed and accuracy for Neural SDEs, outperforming state-of-the-art techniques.
We are interested in strong approximations of one-dimensional SDEs which have non-Lipschitz coefficients and which take values in a domain. Under a set of general assumptions we derive an implicit scheme that preserves the domain of the SDEs and is strongly convergent with rate one. Moreover, we show that this general …
We introduce a mean-reverting SDE whose solution is naturally defined on the space of correlation matrices. This SDE can be seen as an extension of the well-known Wright-Fisher diffusion. We provide conditions that ensure weak and strong uniqueness of the SDE, and describe its ergodic limit. We also shed light on a use…
Neural SDEs reduce variance in stochastic simulations.
problem Efficiency of Monte Carlo simulations in finance.
method Use neural SDEs with control variates parameterized by neural networks.
result Prove optimality conditions for variance reduction in SDEs with infinite activity.
NSFs learn SDE transition laws for efficient sampling.
problem Efficiently sampling between arbitrary time points in SDEs.
method Conditional normalising flows with architectural constraints.
result Up to two orders of magnitude speed-ups at large time gaps.
New schemes for SDEs on manifolds keep solutions close to the manifold.
problem Solving SDEs constrained to manifolds in high accuracy.
method Geometrically invariant numerical schemes that remain close to the manifold.
result The schemes converge under standard assumptions and outperform existing methods.
Generates consistent IV surfaces using VAEs and SDE models.
problem Creating arbitrage-free IV surfaces from historical data.
method Combining VAEs with SDE models for parameter distribution, sampling, and decoding.
result Superior out-of-sample performance of the refined VAE model.
Neural SDEs model continuous sequences using neural networks.
problem Modeling continuous-time dynamics in sequence data.
method Interprets time-series as samples from a continuous dynamical system, parameterized by Neural SDE.
result Demonstrates superior performance in diverse sequence modeling tasks.
TFM trains Neural SDEs without backpropagation, improving clinical time series modeling.
problem Modeling irregularly sampled time series in medicine.
method Trajectory Flow Matching (TFM) using flow matching for generative modeling.
result TFM improves performance on clinical time series datasets.
Deep learning estimates time-varying Markov model parameters.
problem Estimating time-dependent parameters in Markov models.
method Reframes parameter estimation as an optimization problem using maximum likelihood.
result Real solution close to SDE with neural network-derived parameters under specific conditions.
Neural SDEs model suicide risk with compact state space constraints.
problem Modeling suicide risk with irregular, noisy, and partially observed data.
method Developed neural SDEs confined to compact state spaces, addressing domain constraints and numerical stability.
result Improved forecasts and optimization dynamics over standard models on EMA datasets.
SING improves state inference in latent SDE models for better drift function estimation.
problem Intractable posterior inference in latent SDE models.
method Natural gradient variational inference.
result SING provides faster and more reliable inference in latent SDE models.