NSGLD improves SGLD for non-convex optimization problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Boundary rigidity proven for non-reversible Finsler metrics.
DIGing-SGLD improves SGLD for scalable Bayesian learning in dynamic networks.
Characterizes isometries between non-reversible Finsler manifolds.
A new sampler speeds up Bayesian mixture models.
SGLD proves geometric ergodicity via reflection coupling for nonconvex log-concave distributions.
Adaptive algorithm improves convergence rate of Langevin dynamics.
Paper develops robust SGLD for solving non-convex DRO problems.
New bounds for SGLD show error decreases with more data.
We continue our study of geometric analysis on (possibly non-reversible) Finsler manifolds, based on the Bochner inequality established by the author and Sturm. Following the approach of the -calculus a la Bakry et al, we show the dimensional versions of the Poincare--Lichnerowicz inequality, the logarithmic Sobolev…
Paper explores low-precision SGLD for neural networks, reducing costs without sacrificing performance.
RELTA-SGLD stabilizes nonconvex SGLD updates with a lighter taming scheme.
Improved error estimate for SGLD sampling algorithm.
Stochastic Gradient Langevin Dynamics (SGLD) is a popular variant of Stochastic Gradient Descent, where properly scaled isotropic Gaussian noise is added to an unbiased estimate of the gradient at each iteration. This modest change allows SGLD to escape local minima and suffices to guarantee asymptotic convergence to g…
Optimizes SGLD noise structure for better generalization bounds.
As an important Markov Chain Monte Carlo (MCMC) method, stochastic gradient Langevin dynamics (SGLD) algorithm has achieved great success in Bayesian learning and posterior sampling. However, SGLD typically suffers from slow convergence rate due to its large variance caused by the stochastic gradient. In order to allev…
Bayesian deep learning is recently regarded as an intrinsic way to characterize the weight uncertainty of deep neural networks~(DNNs). Stochastic Gradient Langevin Dynamics~(SGLD) is an effective method to enable Bayesian deep learning on large-scale datasets. Previous theoretical studies have shown various appealing p…
Study on how non-reversible diffusion processes affect homology on manifolds.
New approach uses SGLD to minimize CVaR for portfolio weights.
A nonparametric Bayesian sparse graph linear dynamical system (SGLDS) is proposed to model sequentially observed multivariate data. SGLDS uses the Bernoulli-Poisson link together with a gamma process to generate an infinite dimensional sparse random graph to model state transitions. Depending on the sparsity pattern of…
Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally infeasible. The recently proposed stochastic gradient Langevin dynamics (SGLD) method circumvents this problem in three ways: it generates proposed moves using only a subset of the data, it skips the Metropolis-Hastings a…
Novel bounds for SGLD show generalization error decreases with more samples.
New rates for GLD and SGLD in infinite-dimensional spaces without dimensionality issues.
New method improves uncertainty quantification in latent variable models.
Study of parabolas in Funk metric on unit disk.
Effective training of deep neural networks suffers from two main issues. The first is that the parameter spaces of these models exhibit pathological curvature. Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD). These methods improve convergence by adapting to th…
New algorithms for Bayesian inference in decentralized learning.
Stochastic Gradient Langevin Dynamics (SGLD) has emerged as a key MCMC algorithm for Bayesian learning from large scale datasets. While SGLD with decreasing step sizes converges weakly to the posterior distribution, the algorithm is often used with a constant step size in practice and has demonstrated successes in mach…
In the asymmetric setting, Hilbert's fourth problem asks to construct and study all (non-reversible) projective Finsler metrics: Finsler metrics defined on open, convex subsets of real projective -space for which geodesics lie on projective lines. While asymmetric norms and Funk metrics provide many examples of esse…
Langevin dynamics (LD) has been proven to be a powerful technique for optimizing a non-convex objective as an efficient algorithm to find local minima while eventually visiting a global minimum on longer time-scales. LD is based on the first-order Langevin diffusion which is reversible in time. We study two variants th…
Designing deterministic denominators for SGLD stabilizes large drifts.
A new sampler improves the inference of causal structures from observational data.
We develop the basics of a theory of almost isometries for spaces endowed with a quasi-metric. The case of non-reversible Finsler (more specifically, Randers) metrics is of particular interest, and it is studied in more detail. The main motivation arises from General Relativity, and more specifically in spacetimes endo…
The paper analyzes variance reduction in stochastic gradient Langevin dynamics.
Improved convergence for non-log-concave sampling.
Stochastic gradient Langevin dynamics (SGLD) is a fundamental algorithm in stochastic optimization. Recent work by Zhang et al. [2017] presents an analysis for the hitting time of SGLD for the first and second order stationary points. The proof in Zhang et al. [2017] is a two-stage procedure through bounding the Cheege…
In this paper we address the following question: Can we approximately sample from a Bayesian posterior distribution if we are only allowed to touch a small mini-batch of data-items for every sample we generate?. An algorithm based on the Langevin equation with stochastic gradients (SGLD) was previously proposed to solv…
We show that Entropy-SGD (Chaudhari et al., 2017), when viewed as a learning algorithm, optimizes a PAC-Bayes bound on the risk of a Gibbs (posterior) classifier, i.e., a randomized classifier obtained by a risk-sensitive perturbation of the weights of a learned classifier. Entropy-SGD works by optimizing the bound's p…
SGLRW improves robustness of stochastic gradient MCMC methods.
SGLDiff approximates Bayesian posterior distributions with subsampling error.
Paper uses SGLD to recover signals from generative models, proving convergence under mild conditions.
One way to avoid overfitting in machine learning is to use model parameters distributed according to a Bayesian posterior given the data, rather than the maximum likelihood estimator. Stochastic gradient Langevin dynamics (SGLD) is one algorithm to approximate such Bayesian posteriors for large models and datasets. SGL…
The book covers scalable MCMC methods for Bayesian learning.
We show the existence of at least two geometrically distinct closed geodesics on an n-dimensional sphere with a bumpy and non-reversible Finsler metric for n>2.
Bayesian neural networks (BNNs) allow us to reason about uncertainty in a principled way. Stochastic Gradient Langevin Dynamics (SGLD) enables efficient BNN learning by drawing samples from the BNN posterior using mini-batches. However, SGLD and its extensions require storage of many copies of the model parameters, a p…
For non-reversible Finsler metrics of positive flag curvature on spheres and projective spaces we present results about the number and the length of closed geodesics and about their stability properties.
New algorithms improve sampling from Bayesian deep learning models.
DP-BNNs improve accuracy, privacy, and reliability in neural networks.