This work proposes using zero-variance control variates to reduce variance in pathwise gradient estimators for variational inference.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider the task of classification in the high dimensional setting where the number of features of the given data is significantly greater than the number of observations. To accomplish this task, we propose a heuristic, called sparse zero-variance discriminant analysis (SZVD), for simultaneously performing linear …
Develops a method for non-equilibrium importance sampling to estimate expectations and constants.
It is well known that Markov chain Monte Carlo (MCMC) methods scale poorly with dataset size. A popular class of methods for solving this issue is stochastic gradient MCMC. These methods use a noisy estimate of the gradient of the log posterior, which reduces the per iteration computational cost of the algorithm. Despi…
HDT improves MCMC on graphs with history-dependent sampling.
Paper introduces VDE, a variance-reduced determinant estimator.
We examine Kreps' (2019) conjecture that optimal expected utility in the classic Black--Scholes--Merton (BSM) economy is the limit of optimal expected utility for a sequence of discrete-time economies that "approach" the BSM economy in a natural sense: The th discrete-time economy is generated by a scaled -step r…
We develop a scalable method for Bayesian neural networks with stochastic differential equations.
In this note we answer a question of G. Lecué, by showing that column normalization of a random matrix with iid entries need not lead to good sparse recovery properties, even if the generating random variable has a reasonable moment growth. Specifically, for every we construct a random vector …
In recent years, several methods have been proposed for the discovery of causal structure from non-experimental data (Spirtes et al. 2000; Pearl 2000). Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, …
Stochastic gradient descent (SGD) is a popular and efficient method with wide applications in training deep neural nets and other nonconvex models. While the behavior of SGD is well understood in the convex learning setting, the existing theoretical results for SGD applied to nonconvex objective functions are far from …
Study uses multidimensional SE-NBD process to analyze default portfolios and identify shock amplification.
In this paper we consider online mirror descent (OMD) algorithms, a class of scalable online learning algorithms exploiting data geometric structures through mirror maps. Necessary and sufficient conditions are presented in terms of the step size sequence for the convergence of an OMD algorithm with respe…
This study optimizes model averaging for personalized collaborative learning.
We show that -means (Lloyd's algorithm) is obtained as a special case when truncated variational EM approximations are applied to Gaussian Mixture Models (GMM) with isotropic Gaussians. In contrast to the standard way to relate -means and GMMs, the provided derivation shows that it is not required to consider Gau…
We explore a new research direction in Bayesian variational inference with discrete latent variable priors where we exploit Kronecker matrix algebra for efficient and exact computations of the evidence lower bound (ELBO). The proposed "DIRECT" approach has several advantages over its predecessors; (i) it can exactly co…
We propose a general yet simple theorem describing the convergence of SGD under the arbitrary sampling paradigm. Our theorem describes the convergence of an infinite array of variants of SGD, each of which is associated with a specific probability law governing the data selection rule used to form mini-batches. This is…
We develop importance sampling based efficient simulation techniques for three commonly encountered rare event probabilities associated with random walks having i.i.d. regularly varying increments; namely, 1) the large deviation probabilities, 2) the level crossing probabilities, and 3) the level crossing probabilities…
This work improves distribution recovery from sparse data using Random Forest implicit regularization.
Optimization can learn Johnson-Lindenstrauss embeddings without randomization.
SoftCVI uses contrastive estimation to infer complex posteriors.
Extends linear structural causal models to include deterministic relations and latent confounders for causal discovery.
HAMD optimizes cubic portfolios without quadratization, achieving better results.
SRMC framework reduces Monte Carlo variance by history-based sampling in high-dimensional spaces.
Improves inference-time alignment for diffusion models without updating weights.