Nonlinear MCMC improves Bayesian machine learning sampling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New analysis of SGD with MCMC gradient estimator shows convergence rate and saddle point escape.
Cyclical MCMC tackles high-dimensional multimodal distributions, showing convergence under certain conditions.
This paper studies a curious phenomenon in learning energy-based model (EBM) using MCMC. In each learning iteration, we generate synthesized examples by running a non-convergent, non-mixing, and non-persistent short-run MCMC toward the current model, always starting from the same initial distribution such as uniform no…
Recent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gradient thermostat. While finite-time convergence properties of th…
It is known that the Langevin dynamics used in MCMC is the gradient flow of the KL divergence on the Wasserstein space, which helps convergence analysis and inspires recent particle-based variational inference methods (ParVIs). But no more MCMC dynamics is understood in this way. In this work, by developing novel conce…
This review article surveys data augmentation MCMC algorithms.
Stochastic EM with biased MCMC improves inference stability.
Stochastic gradient Markov Chain Monte Carlo (SG-MCMC) has been developed as a flexible family of scalable Bayesian sampling algorithms. However, there has been little theoretical analysis of the impact of minibatch size to the algorithm's convergence rate. In this paper, we prove that under a limited computational bud…
This study compares parallel SMC and MCMC for Bayesian deep learning, showing SMC parallel is faster.
Normalized random measures (NRMs) provide a broad class of discrete random measures that are often used as priors for Bayesian nonparametric models. Dirichlet process is a well-known example of NRMs. Most of posterior inference methods for NRM mixture models rely on MCMC methods since they are easy to implement and the…
Stochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based on some outdated par…
This paper analyzes MCMC algorithms on large graphs using Dirichlet forms.
New method reduces computational cost for Bayesian inference.
New couplings improve understanding of molecular dynamics convergence.
LSB is a new MCMC method for discrete spaces that reduces target evaluations.
This paper develops tools for nonreversible MCMC with convergence guarantees.
Markov Chain Monte Carlo (MCMC) methods such as Gibbs sampling are finding widespread use in applied statistics and machine learning. These often lead to difficult computational problems, which are increasingly being solved on parallel and distributed systems such as compute clusters. Recent work has proposed running i…
AMAGOLD improves stochastic gradient MCMC by infrequent Metropolis-Hastings corrections.
AI-driven framework optimizes MCMC-based preconditioners for faster linear system solving.
Markov chain Monte Carlo (MCMC) methods have not been broadly adopted in Bayesian neural networks (BNNs). This paper initially reviews the main challenges in sampling from the parameter posterior of a neural network via MCMC. Such challenges culminate to lack of convergence to the parameter posterior. Nevertheless, thi…
Statistical inference methods are fundamentally important in machine learning. Most state-of-the-art inference algorithms are variants of Markov chain Monte Carlo (MCMC) or variational inference (VI). However, both methods struggle with limitations in practice: MCMC methods can be computationally demanding; VI methods …
New variational flows improve Monte Carlo and normalization tasks.
The book covers scalable MCMC methods for Bayesian learning.
Markov chain Monte Carlo (MCMC) is widely regarded as one of the most important algorithms of the 20th century. Its guarantees of asymptotic convergence, stability, and estimator-variance bounds using only unnormalized probability functions make it indispensable to probabilistic programming. In this paper, we introduce…
The posteriors over neural network weights are high dimensional and multimodal. Each mode typically characterizes a meaningfully different representation of the data. We develop Cyclical Stochastic Gradient MCMC (SG-MCMC) to automatically explore such distributions. In particular, we propose a cyclical stepsize schedul…
This project compares MCMC and VI for Bayesian PMF on MovieLens.
New KSDs control moments in approximations, improving diagnostics and tests.
New MCMC methods map high-dimensional problems to spheres for better mixing.
Acyclic digraphs are the underlying representation of Bayesian networks, a widely used class of probabilistic graphical models. Learning the underlying graph from data is a way of gaining insights about the structural properties of a domain. Structure learning forms one of the inference challenges of statistical graphi…
Improved sampling accuracy in SG-MCMC methods via non-uniform gradient subsampling.
New Langevin method achieves third order convergence for strongly log-concave distributions.
New algorithm MTMC reduces MCMC evaluation costs.
Deep unfolding accelerates MCMC-based COP solvers.
Recent works have derived non-asymptotic upper bounds for convergence of underdamped Langevin MCMC. We revisit these bound and consider introducing scaling terms in the underlying underdamped Langevin equation. In particular, we provide conditions under which an appropriate scaling allows to improve the error bounds in…
New model improves MCMC efficiency and multi-modal distribution exploration.
This paper tackles sampling issues in latent space EBMs by introducing diffusion-based amortization.
Stochastic Stein Discrepancies improve inference efficiency.
Split-Merge MCMC (Monte Carlo Markov Chain) is one of the essential and popular variants of MCMC for problems when an MCMC state consists of an unknown number of components. It is well known that state-of-the-art methods for split-merge MCMC do not scale well. Strategies for rapid mixing requires smart and informative …
Optimizes MCMC chains with neural control variates.
Neural network MCMC sampler maximizes proposal entropy for efficient sampling.
Bayesian neural networks tutorial via MCMC in Python.
SMTM improves MCMC sampling in high dimensions with multiple proposals and stereographic integration.
Variational inference lies at the core of many state-of-the-art algorithms. To improve the approximation of the posterior beyond parametric families, it was proposed to include MCMC steps into the variational lower bound. In this work we explore this idea using steps of the Hamiltonian Monte Carlo (HMC) algorithm, an e…
We propose a new class of learning algorithms that combines variational approximation and Markov chain Monte Carlo (MCMC) simulation. Naive algorithms that use the variational approximation as proposal distribution can perform poorly because this approximation tends to underestimate the true variance and other features…
Improved sampling for Bayesian neural networks reduces vanishing acceptance rates and increases predictive accuracy.
KSD Thinning uses KSD to thin MCMC samples efficiently.
Study improves Bayesian calibration of mechanical properties using active learning and MCMC.