A new metric tensor improves Riemann manifold Monte Carlo for Bayesian models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Log-density gradient estimation is a fundamental statistical problem and possesses various practical applications such as clustering and measuring non-Gaussianity. A naive two-step approach of first estimating the density and then taking its log-gradient is unreliable because an accurate density estimate does not neces…
Proposes log density gradient to improve reinforcement learning sample complexity.
Improved VI with Price's gradient estimator for target log-density.
Pathfinder uses quasi-Newton optimization for variational inference.
Modal regression is aimed at estimating the global mode (i.e., global maximum) of the conditional density function of the output variable given input variables, and has led to regression methods robust against heavy-tailed or skewed noises. The conditional mode is often estimated through maximization of the modal regre…
Non-Gaussian component analysis (NGCA) is aimed at identifying a linear subspace such that the projected data follows a non-Gaussian distribution. In this paper, we propose a novel NGCA algorithm based on log-density gradient estimation. Unlike existing methods, the proposed NGCA algorithm identifies the linear subspac…
Estimates Gaussian location model with ridge regularization, comparing variational and spectral methods.
A new sampling method, RC-LMC, reduces computational cost for high-dimensional log-concave distributions.
Smart Bayes integrates generative and discriminative features for improved classification.
We consider the problem of sampling from a strongly log-concave density in , and prove an information theoretic lower bound on the number of stochastic gradient queries of the log density needed. Several popular sampling algorithms (including many Markov chain Monte Carlo methods) operate by using stochas…
Non-Gaussian component analysis (NGCA) is an unsupervised linear dimension reduction method that extracts low-dimensional non-Gaussian "signals" from high-dimensional data contaminated with Gaussian noise. NGCA can be regarded as a generalization of projection pursuit (PP) and independent component analysis (ICA) to mu…
Mean shift clustering finds the modes of the data probability density by identifying the zero points of the density gradient. Since it does not require to fix the number of clusters in advance, the mean shift has been a popular clustering algorithm in various application fields. A typical implementation of the mean shi…
BBVI converges nearly dimensionally independent for log-concave targets.
Langevin Monte Carlo (LMC) is an iterative algorithm used to generate samples from a distribution that is known only up to a normalizing constant. The nonasymptotic dependence of its mixing time on the dimension and target accuracy is understood mainly in the setting of smooth (gradient-Lipschitz) log-densities, a seri…
Noise-corrected Langevin algorithm improves sampling from noisy data.
In this paper, we study the problem of sampling from a given probability density function that is known to be smooth and strongly log-concave. We analyze several methods of approximate sampling based on discretizations of the (highly overdamped) Langevin diffusion and establish guarantees on its error measured in the W…
Novel criterion identifies heteroscedastic noise in causal discovery.
Flow-based generative models parameterize probability distributions through an invertible transformation and can be trained by maximum likelihood. Invertible residual networks provide a flexible family of transformations where only Lipschitz conditions rather than strict architectural constraints are needed for enforci…
ASVGD accelerates SVGD for efficient sampling from Gaussian targets.
Parallel sampling for smooth distributions with fast convergence.
AR-DAE approximates entropy gradient for machine learning models.
Proximal Diffusion Models improve generative model efficiency.
ASVGD accelerates SVGD for efficient sampling.
DPS uses PINNs to estimate drift in diffusion models for sampling.
Improved convergence for non-log-concave sampling.
In this paper, we provide new insights on the Unadjusted Langevin Algorithm. We show that this method can be formulated as a first order optimization algorithm of an objective functional defined on the Wasserstein space of order . Using this interpretation and techniques borrowed from convex optimization, we give a …
We study the problem of sampling from a probability distribution on $\rset^d$ which has a density \wrt\ the Lebesgue measure known up to a normalization factor $x \mapsto \rme^{-U(x)} / \int_{\rset^d} \rme^{-U(y)} \rmd y$. We analyze a sampling method based on the Euler discretization of the Langevin stochastic dif…
We establish general conditions under which Markov chains produced by the Hamiltonian Monte Carlo method will and will not be geometrically ergodic. We consider implementations with both position-independent and position-dependent integration times. In the former case we find that the conditions for geometric ergodicit…
Enhances normal mean estimation with side info using NIT approach.
We present a new method for evaluating and training unnormalized density models. Our approach only requires access to the gradient of the unnormalized model's log-density. We estimate the Stein discrepancy between the data density and the model density defined by a vector function of the data. We paramete…
Proposes a deep neural network for multi-dimensional functional data classification.
New method for estimating diffusion model densities without solving flows.
New method improves online covariance estimation for SGD.
The kernel exponential family is a rich class of distributions, which can be fit efficiently and with statistical guarantees by score matching. Being required to choose a priori a simple kernel such as the Gaussian, however, limits its practical applicability. We provide a scheme for learning a kernel parameterized by …
The gradient noise of SGD is considered to play a central role in the observed strong generalization abilities of deep learning. While past studies confirm that the magnitude and the covariance structure of gradient noise are critical for regularization, it remains unclear whether or not the class of noise distribution…
Normalizing flow regression approximates posterior distributions without additional sampling.
The paper optimizes regret using covariance between costs and decisions.
GBMixed boosts mixed models for clustered data, estimating mean and variance flexibly.
Improved score matching methods for estimating score functions and Hessians without high dimensionality.
New method trains EBMs using NFs for more accurate likelihood estimation.
Unified view of score estimators for flexible densities.
Density estimation is a fundamental problem in statistical learning. This problem is especially challenging for complex high-dimensional data due to the curse of dimensionality. A promising solution to this problem is given here in an inference-free hierarchical framework that is built on score matching. We revisit the…
A new variational inference method using Gaussian score matching.
In applications of Gaussian processes where quantification of uncertainty is of primary interest, it is necessary to accurately characterize the posterior distribution over covariance parameters. This paper proposes an adaptation of the Stochastic Gradient Langevin Dynamics algorithm to draw samples from the posterior …
Statistical analysis of algorithm unrolling for inverse problems.
Gradient flow solves optimal mass transport for covariance matrices.
MonoFlow rethinks GANs using Wasserstein gradient flows.