A new algorithm FastGM speeds up generating Gumbel-Max variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a new estimator for generic discrete distributions.
Paper improves Gumbel-Softmax estimator variance reduction.
Categorical variables are a natural choice for representing discrete structure in the world. However, stochastic neural networks rarely use categorical latent variables due to the inability to backpropagate through samples. In this work, we present an efficient gradient estimator that replaces the non-differentiable sa…
The Gumbel trick is a method to sample from a discrete probability distribution, or to estimate its normalizing partition function. The method relies on repeatedly applying a random perturbation to the distribution in a particular way, each time solving for the most likely configuration. We derive an entire family of r…
Boltzmann machines (BMs) are appealing candidates for powerful priors in variational autoencoders (VAEs), as they are capable of capturing nontrivial and multi-modal distributions over discrete variables. However, non-differentiability of the discrete units prohibits using the reparameterization trick, essential for lo…
Reparameterization of variational auto-encoders with continuous random variables is an effective method for reducing the variance of their gradient estimates. In the discrete case, one can perform reparametrization using the Gumbel-Max trick, but the resulting objective relies on an operation and is non-dif…
Unified framework for gradient estimation in combinatorial spaces.
Permutations and matchings are core building blocks in a variety of latent variable models, as they allow us to align, canonicalize, and sort data. Learning in such models is difficult, however, because exact marginalization over these combinatorial objects is intractable. In response, this paper introduces a collectio…
Unified approach to DP problems using Gumbel distribution and variational Bayesian inference.
This article proposes a method to quantify the structure of a bipartite graph using a network entropy per link. The network entropy of a bipartite graph with random links is calculated both numerically and theoretically. As an application of the proposed method to analyze collective behavior, the affairs in which parti…
The tail of the distribution of a sum of a random number of independent and identically distributed nonnegative random variables depends on the tails of the number of terms and of the terms themselves. This situation is of interest in the collective risk model, where the total claim size in a portfolio is the sum of a …
GDM models time series with smoother transitions and interpretable states.
We present a probabilistic framework for studying adversarial attacks on discrete data. Based on this framework, we derive a perturbation-based method, Greedy Attack, and a scalable learning-based method, Gumbel Attack, that illustrate various tradeoffs in the design of attacks. We demonstrate the effectiveness of thes…
Improved Gumbel watermark detection method.
Inference in log-linear models scales linearly in the size of output space in the worst-case. This is often a bottleneck in natural language processing and computer vision tasks when the output space is feasibly enumerable but very large. We propose a method to perform inference in log-linear models with sublinear amor…
Many problems in real life can be converted to combinatorial optimization problems (COPs) on graphs, that is to find a best node state configuration or a network structure such that the designed objective function is optimized under some constraints. However, these problems are notorious for their hardness to solve bec…
The Gumbel-max trick and its extensions simplify sampling from categorical distributions in machine learning.
Thompson sampling has impressive empirical performance for many multi-armed bandit problems. But current algorithms for Thompson sampling only work for the case of conjugate priors since these algorithms require to infer the posterior, which is often computationally intractable when the prior is not conjugate. In this …
Modified EAT method improves Poisson gradient estimation.
A new algorithm optimizes graph problems faster and more accurately.
Develops a framework to test excessive influence of small data subsets.
Proposes a new method for estimating counterfactual treatment effects.
Generative Adversarial Networks (GAN) have limitations when the goal is to generate sequences of discrete elements. The reason for this is that samples from a distribution on discrete objects such as the multinomial are not differentiable with respect to the distribution parameters. This problem can be avoided by using…
Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult. We propose a new class of smoothing transformations based on a mixture of two overlapping distributions, and show that the proposed transformation can be used for training binary lat…
The problem of drawing samples from a discrete distribution can be converted into a discrete optimization problem. In this work, we show how sampling from a continuous distribution can be converted into an optimization problem over continuous space. Central to the method is a stochastic process recently described in ma…
Paper proposes a new method for predicting drug interactions using adversarial autoencoders.
Investigates statistical properties of perturb-softmax and perturb-argmax distributions.
This paper is dedicated to the consistency of systemic risk measures with respect to stochastic dependence. It compares two alternative notions of Conditional Value-at-Risk (CoVaR) available in the current literature. These notions are both based on the conditional distribution of a random variable Y given a stress eve…
Heterogeneity of economic agents is emphasized in a new trend of macroeconomics. Accordingly the new emerging discipline requires one to replace the production function, one of key ideas in the conventional economics, by an alternative which can take an explicit account of distribution of firms' production activities. …
The well-known Gumbel-Max trick for sampling from a categorical distribution can be extended to sample elements without replacement. We show how to implicitly apply this 'Gumbel-Top-' trick on a factorized distribution over sequences, allowing to draw exact samples without replacement using a Stochastic Beam Sea…
The Gumbel-Softmax is a continuous distribution over the simplex that is often used as a relaxation of discrete distributions. Because it can be readily interpreted and easily reparameterized, it enjoys widespread use. We propose a modular and more flexible family of reparameterizable distributions where Gaussian noise…
Study on estimating Gumbel--Max watermark proportions in edited documents.
This paper proposes a novel learning method for multi-task applications. Multi-task neural networks can learn to transfer knowledge across different tasks by using parameter sharing. However, sharing parameters between unrelated tasks can hurt performance. To address this issue, we propose a framework to learn fine-gra…
In many applications we seek to maximize an expectation with respect to a distribution over discrete variables. Estimating gradients of such objectives with respect to the distribution parameters is a challenging problem. We analyze existing solutions including finite-difference (FD) estimators and continuous relaxatio…
New method for efficient marginalization of discrete latent variables in neural networks.
The paper introduces a method to improve adversarial robustness in neural networks using randomized perturbations.
Improved learning of probabilistic box embeddings by modeling parameters with Gumbel distributions.
A new method for categorical variational inference using discrete normalizing flows.
Improved CAEs reduce training time and enhance generalization.
We introduce a new stochastic smoothing perspective to study adversarial contextual bandit problems. We propose a general algorithm template that represents random perturbation based algorithms and identify several perturbation distributions that lead to strong regret bounds. Using the idea of smoothness, we provide an…
For network architecture search (NAS), it is crucial but challenging to simultaneously guarantee both effectiveness and efficiency. Towards achieving this goal, we develop a differentiable NAS solution, where the search space includes arbitrary feed-forward network consisting of the predefined number of connections. Be…
Neural jump model improves option pricing accuracy.
This paper improves MADDPG's performance in discrete grid-world scenarios.
FAIR-NN finds invariant variables for causal inference across diverse environments.
Deep learning improves community detection in graph datasets.
Estimates proportions of LLM-generated text in mixed documents.
Paper compares econometric models with machine learning for energy forecasting.