Boltzmann machines are powerful distributions that have been shown to be an effective prior over binary latent variables in variational autoencoders (VAEs). However, previous methods for training discrete VAEs have used the evidence lower bound and not the tighter importance-weighted bound. We propose two approaches fo…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper develops sum-of-squares relaxations for computing -divergences.
A new method for categorical variational inference using discrete normalizing flows.
Paper improves variational inference on Boolean hypercube using quantum methods.
We consider a relaxed notion of energy of non-parametric codimension one surfaces that takes account of area, mean curvature, and Gauss curvature. It is given by the best value obtained by approximation with inscribed polyhedral surfaces. The BV and measure properties of functions with finite relaxed energy are studied…
This work interprets SFA through variational inference, relaxing linearity constraints.
Learning in models with discrete latent variables is challenging due to high variance gradient estimators. Generally, approaches have relied on control variates to reduce the variance of the REINFORCE estimator. Recent work (Jang et al. 2016, Maddison et al. 2016) has taken a different approach, introducing a continuou…
Due to challenging applications such as collaborative filtering, the matrix completion problem has been widely studied in the past few years. Different approaches rely on different structure assumptions on the matrix in hand. Here, we focus on the completion of a (possibly) low-rank matrix with binary entries, the so-c…
The paper studies harmonic maps to the circle with complex singular sets.
New method relaxes TV distance for two-sample testing without distributional assumptions.
Estimate relaxation times in nonextensive systems using gradient flow for Tsallis entropy maximization.
Defines weak normals for irregular curves in high-dimensional spaces.
New method improves GP regression by relaxing variational assumption.
CO-BED optimizes experiments using Bayesian methods and information theory.
Regularized regression problems are ubiquitous in statistical modeling, signal processing, and machine learning. Sparse regression in particular has been instrumental in scientific model discovery, including compressed sensing applications, variable selection, and high-dimensional analysis. We propose a broad framework…
Improved hierarchical discrete VAEs for better stability and performance.
We introduce a globally-convergent algorithm for optimizing the tree-reweighted (TRW) variational objective over the marginal polytope. The algorithm is based on the conditional gradient method (Frank-Wolfe) and moves pseudomarginals within the marginal polytope through repeated maximum a posteriori (MAP) calls. This m…
Let α(s) be an arc on a connected oriented surface S in E3, parameterized by arc length s, with torsion τ and length l. The total square torsion F of α is defined by T=\int_{0}^{l}τ^{2}ds\ $. . The arc α is called a relaxed elastic line of second kind if it is an extremal for the variational problem of minimizing the v…
Let be an arc on a connected oriented surface in Minkowski 3-space, parameterized by arc length , with torsion and length . The total square torsion of is defined by . The arc is called a relaxed elastic line of second kind if it is an extremal for the variational prob…
Paper relaxes differential privacy for correlated features, improving privacy-utility trade-off.
In this paper, we consider the classical variational problem in the Galilean space. we develop the Euler-Lagrange equations for a elastic line on an oriented surface in the Galilean 3-dimensional space . Using the varia- tion method, we will try to give some characterization for the solution curve (the elastic lin…
Optimal neural network approximation for Wasserstein gradient direction via convex optimization.
Differentiable relaxation for inferring partial orders from noisy linear data.
DisARM improves gradient estimation for binary latent variables.
DD-VAE uses deterministic decoding for better latent code utilization in discrete data.
Generative models of graphs are well-known, but many existing models are limited in scalability and expressivity. We present a novel sequential graphical variational autoencoder operating directly on graphical representations of data. In our model, the encoding and decoding of a graph as is framed as a sequential decon…
Sequential coordinate ascent is more robust in high-dimensional linear regression.
In many applications we seek to maximize an expectation with respect to a distribution over discrete variables. Estimating gradients of such objectives with respect to the distribution parameters is a challenging problem. We analyze existing solutions including finite-difference (FD) estimators and continuous relaxatio…
Beta process is the standard nonparametric Bayesian prior for latent factor model. In this paper, we derive a structured mean-field variational inference algorithm for a beta process non-negative matrix factorization (NMF) model with Poisson likelihood. Unlike the linear Gaussian model, which is well-studied in the non…
We present a framework for learning disentangled and interpretable jointly continuous and discrete representations in an unsupervised manner. By augmenting the continuous latent distribution of variational autoencoders with a relaxed discrete distribution and controlling the amount of information encoded in each latent…
A new generative model relaxes the bijectivity requirement for invertible flows.
A key challenge for gradient based optimization methods in model-free reinforcement learning is to develop an approach that is sample efficient and has low variance. In this work, we apply Kronecker-factored curvature estimation technique (KFAC) to a recently proposed gradient estimator for control variate optimization…
VaSST uses soft symbolic trees for probabilistic symbolic regression.
PIVID infers DAG structures from data using variational inference and permutations.
Reparameterization of variational auto-encoders with continuous random variables is an effective method for reducing the variance of their gradient estimates. In the discrete case, one can perform reparametrization using the Gumbel-Max trick, but the resulting objective relies on an operation and is non-dif…
Using the theory of group action, we first introduce the concept of the automorphism group of an exponential family or a graphical model, thus formalizing the general notion of symmetry of a probabilistic model. This automorphism group provides a precise mathematical framework for lifted inference in the general expone…
Unified framework for fair regression in aware and unaware settings.
Two non-local asymptotic invariants of magnetic fields for the ideal magnetohydrodynamics are introduced. The velocity of variation of the invariants for a non-ideal magnetohydrodynamics with a small magnetic dissipation is estimated. By means of the invariants the spectra of electromagnetic fields are investigated. A …
Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult. We propose a new class of smoothing transformations based on a mixture of two overlapping distributions, and show that the proposed transformation can be used for training binary lat…
Paper develops efficient variational inference for sparse deep learning with theoretical guarantees.
VSI model predicts survival distributions efficiently.
GCVAE improves disentanglement in VAEs while balancing reconstruction error.
We derive streamlined mean field variational Bayes algorithms for fitting linear mixed models with crossed random effects. In the most general situation, where the dimensions of the crossed groups are arbitrarily large, streamlining is hindered by lack of sparseness in the underlying least squares system. Because of th…
Paper tightens variational GP approximations for large datasets.
Practitioners of Bayesian statistics have long depended on Markov chain Monte Carlo (MCMC) to obtain samples from intractable posterior distributions. Unfortunately, MCMC algorithms are typically serial, and do not scale to the large datasets typical of modern machine learning. The recently proposed consensus Monte Car…
Proposes a robust VIB approach using soft labels and mutual info estimation.
A new framework for sparse regression models with slow variations.
Modern neural network training relies on piece-wise (sub-)differentiable functions in order to use backpropagation to update model parameters. In this work, we introduce a novel method to allow simple non-differentiable functions at intermediary layers of deep neural networks. We do so by training with a differentiable…