Stochastic variational inference (SVI) plays a key role in Bayesian deep learning. Recently various divergences have been proposed to design the surrogate loss for variational inference. We present a simple upper bound of the evidence as the surrogate loss. This evidence upper bound (EUBO) equals to the log marginal li…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper establishes lower bounds for non-stationary kernelized bandits.
Variational Inference is a powerful tool in the Bayesian modeling toolkit, however, its effectiveness is determined by the expressivity of the utilized variational distributions in terms of their ability to match the true posterior distribution. In turn, the expressivity of the variational family is largely limited by …
The paper analyzes variational autoencoders for state space models with risk bounds.
We consider a non-stationary sequential stochastic optimization problem, in which the underlying cost functions change over time under a variation budget constraint. We propose an -variation functional to quantify the change, which yields less variation for dynamic function sequences whose changes are constrai…
The paper develops estimators for variance in graph structures using fused lasso.
Non-negative matrix factorization (NMF) is a knowledge discovery method that is used in many fields. Variational inference and Gibbs sampling methods for it are also wellknown. However, the variational approximation error has not been clarified yet, because NMF is not statistically regular and the prior distribution us…
Variational inference (VI) is widely used as an efficient alternative to Markov chain Monte Carlo. It posits a family of approximating distributions and finds the closest member to the exact posterior . Closeness is usually measured via a divergence from to . While successful, this approach al…
The total variation distance is a core statistical distance between probability measures that satisfies the metric axioms, with value always falling in . This distance plays a fundamental role in machine learning and signal processing: It is a member of the broader class of -divergences, and it is related to …
This work provides statistical guarantees for VAEs using PAC-Bayesian theory.
We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an algorithm and provide performance guarantees for the regret evaluated against the…
We obtain upper bounds for the eigenvalues of the Schrödinger operator depending on integral quantities of the potential and a conformal invariant called the min-conformal volume. Moreover, when the Schrödinger operator is positive, integral quantities of which appear in upper bounds, can be repla…
In this article we study the second variation of the energy functional associated to the Allen-Cahn equation on closed manifolds. Extending well known analogies between the gradient theory of phase transitions and the theory of minimal hypersurfaces, we prove the upper semicontinuity of the eigenvalues of the stability…
We introduce the Variational Holder (VH) bound as an alternative to Variational Bayes (VB) for approximate Bayesian inference. Unlike VB which typically involves maximization of a non-convex lower bound with respect to the variational parameters, the VH bound involves minimization of a convex upper bound to the intract…
uHMC achieves fast mixing in high dimensions with gradient evaluations.
Sharp bounds on neural network approximation rates and widths.
We use variational methods and a modified curvature flow to give an alternative proof of the existence of a self-shrinking torus under mean curvature flow. As a consequence of the proof, we establish an upper bound for the weighted energy of our shrinking doughnuts.
Semi-implicit variational inference (SIVI) is introduced to expand the commonly used analytic variational distribution family, by mixing the variational parameter with a flexible distribution. This mixing distribution can assume any density function, explicit or not, as long as independent random samples can be generat…
Paper develops a new RL method for MDPs with uncertainty, achieving better regret bounds.
New method for tensor completion using nonconvex dual total variation.
New method relaxes TV distance for two-sample testing without distributional assumptions.
We focus on the maximum regularization parameter for anisotropic total-variation denoising. It corresponds to the minimum value of the regularization parameter above which the solution remains constant. While this value is well know for the Lasso, such a critical value has not been investigated in details for the total…
We analyze variational inference for highly symmetric graphical models such as those arising from first-order probabilistic models. We first show that for these graphical models, the tree-reweighted variational objective lends itself to a compact lifted formulation which can be solved much more efficiently than the sta…
This paper introduces a method to estimate log-likelihood in VAE models.
In this note, we study the relationship between the variational gap and the variance of the (log) likelihood ratio. We show that the gap can be upper bounded by some form of dispersion measure of the likelihood ratio, which suggests the bias of variational inference can be reduced by making the distribution of the like…
Paper improves variational inference on Boolean hypercube using quantum methods.
Proposes CLUB for reliable MI minimization in high dimensions.
New algorithm uses control variates to improve multi-armed bandit performance.
Upper bound on index of rotationally symmetric self-shrinking tori.
New bounds derived using conditional -information for machine learning models.
Recent research has made significant progress on the problem of bounding log partition functions for exponential family graphical models. Such bounds have associated dual parameters that are often used as heuristic estimates of the marginal probabilities required in inference and learning. However these variational est…
For a risk vector , whose components are shared among agents by some random mechanism, we obtain asymptotic lower and upper bounds for the individual agents' exposure risk and the aggregated risk in the market. Risk is measured by Value-at-Risk or Conditional Tail Expectation. We assume Pareto tails for the componen…
Paper bridges VAEs and KDEs for more flexible posterior estimation.
Through using the semidiameter (in connection to: the mean radius and surface radius) of a convex closed hypersurface in as an sharp upper bound of the variational -capacity radius, this paper settles a restriction/variant of S.-T. Yau's \cite[Problem 59]{Yau} from the surface area to t…
We propose a general variational framework of fair clustering, which integrates an original Kullback-Leibler (KL) fairness term with a large class of clustering objectives, including prototype or graph based. Fundamentally different from the existing combinatorial and spectral solutions, our variational multi-term appr…
Kernel SIVI improves variational inference by avoiding lower-level optimization.
Optimal pre-processing reduces disparate impact by minimizing total variation distance.
Computing the partition function of a discrete graphical model is a fundamental inference challenge. Since this is computationally intractable, variational approximations are often used in practice. Recently, so-called gauge transformations were used to improve variational lower bounds on . In this paper, we pro…
The paper tackles approximate unlearning from a subset of training data using variational inference.
Automating statistical modelling is a challenging problem in artificial intelligence. The Automatic Statistician takes a first step in this direction, by employing a kernel search algorithm with Gaussian Processes (GP) to provide interpretable statistical models for regression problems. However this does not scale due …
Many problems in machine learning are naturally expressed in the language of undirected graphical models. Here, we propose black-box learning and inference algorithms for undirected models that optimize a variational approximation to the log-likelihood of the model. Central to our approach is an upper bound on the log-…
The paper improves Gaussian process regression by optimizing hyperparameters.
Study clusters distributions with known or unknown clusters using distribution testing.
New algorithm tackles non-stationary RL with near-optimal regret bounds.
We introduce a variational framework to learn the activation functions of deep neural networks. Our aim is to increase the capacity of the network while controlling an upper-bound of the actual Lipschitz constant of the input-output relation. To that end, we first establish a global bound for the Lipschitz constant of …
Variational Optimization forms a differentiable upper bound on an objective. We show that approaches such as Natural Evolution Strategies and Gaussian Perturbation, are special cases of Variational Optimization in which the expectations are approximated by Gaussian sampling. These approaches are of particular interest …
Marginal MAP problems are notoriously difficult tasks for graphical models. We derive a general variational framework for solving marginal MAP problems, in which we apply analogues of the Bethe, tree-reweighted, and mean field approximations. We then derive a "mixed" message passing algorithm and a convergent alternati…
Sharp inequality between TV and Hellinger distances for Gaussian mixtures.