This paper simplifies ANS for statisticians, making it easier to use.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generalizes bits back coding for time-series models with latent Markov structures.
The bits-back argument suggests that latent variable models can be turned into lossless compression schemes. Translating the bits-back argument into efficient and practical lossless compression schemes for general latent variable models, however, is still an open problem. Bits-Back with Asymmetric Numeral Systems (BB-A…
Deep latent variable models have seen recent success in many data domains. Lossless compression is an application of these models which, despite having the potential to be highly useful, has yet to be implemented in a practical manner. We present `Bits Back with ANS' (BB-ANS), a scheme to perform lossless compression w…
New method improves image compression using bits-back coding.
A new method, REC, compresses images by encoding their latent representations efficiently.
SHVC improves image compression with fewer parameters.
Likelihood-based generative models are the backbones of lossless compression due to the guaranteed existence of codes with lengths close to negative log likelihood. However, there is no guaranteed existence of computationally efficient codes that achieve these lengths, and coding algorithms must be hand-tailored to spe…
Develops a method for lossless compression using latent variable models.
While deep neural networks are a highly successful model class, their large memory footprint puts considerable strain on energy consumption, communication bandwidth, and storage requirements. Consequently, model size reduction has become an utmost goal in deep learning. A typical approach is to train a set of determini…
Study of skateboard flips as continuous curves in group.
Nash's theorem proved with Günther's trick
Explains Conway's tangle trick and its mathematical origins.
Unified framework for gradient estimation in combinatorial spaces.
We make the following striking observation: fully convolutional VAE models trained on 32x32 ImageNet can generalize well, not just to 64x64 but also to far larger photographs, with no changes to the model. We use this property, applying fully convolutional models to lossless compression, demonstrating a method to scale…
We prove all knots can be transformed into a trefoil using special diagrams.
Geometric trick simplifies link homotopy and concordance.
Outlier based Robust Principal Component Analysis (RPCA) requires centering of the non-outliers. We show a "bias trick" that automatically centers these non-outliers. Using this bias trick we obtain the first RPCA algorithm that is optimal with respect to centering.
The Gumbel-max trick and its extensions simplify sampling from categorical distributions in machine learning.
A new gradient estimator for categorical distributions reduces bias and variance.
Establishes necessary and sufficient conditions for smooth triviality of Lie subalgebras and Lie ideals, and proves Moser's trick for foliations.
Retail Product Image Classification is an important Computer Vision and Machine Learning problem for building real world systems like self-checkout stores and automated retail execution evaluation. In this work, we present various tricks to increase accuracy of Deep Learning models on different types of retail product …
Triple-point Whitney trick classifies ornaments of 3-manifolds.
Alexander trick applied to homology spheres for manifold homeomorphisms.
Embolic volume of compact manifolds is defined in terms of Berger's embolic inequality. In this paper, we show a result of relating embolic volume to the first Betti number. The proof relies on Gromov's covering argument appeared in systolic geometry. Berger called this method covering trick. We exploit and present mor…
Expands Bredon's trick for applications in geometry and topology.
Bredon's trick helps extend local properties to global topological spaces.
We observe that gradients computed via the reparameterization trick are in direct correspondence with solutions of the transport equation in the formalism of optimal transport. We use this perspective to compute (approximate) pathwise gradients for probability distributions not directly amenable to the reparameterizati…
We introduce a family of pairwise stochastic gradient estimators for gradients of expectations, which are related to the log-derivative trick, but involve pairwise interactions between samples. The simplest example of our new estimator, dubbed the fundamental trick estimator, is shown to arise from either a) introducin…
The Gumbel trick is a method to sample from a discrete probability distribution, or to estimate its normalizing partition function. The method relies on repeatedly applying a random perturbation to the distribution in a particular way, each time solving for the most likely configuration. We derive an entire family of r…
An online reinforcement learning algorithm is anytime if it does not need to know in advance the horizon T of the experiment. A well-known technique to obtain an anytime algorithm from any non-anytime algorithm is the "Doubling Trick". In the context of adversarial or stochastic multi-armed bandits, the performance of …
Inference in popular nonparametric Bayesian models typically relies on sampling or other approximations. This paper presents a general methodology for constructing novel tractable nonparametric Bayesian methods by applying the kernel trick to inference in a parametric Bayesian model. For example, Gaussian process regre…
New trick builds hyperbolic manifolds from compact ones, proving some don't virtually fiber.
The reparameterization trick has become one of the most useful tools in the field of variational inference. However, the reparameterization trick is based on the standardization transformation which restricts the scope of application of this method to distributions that have tractable inverse cumulative distribution fu…
Boltzmann machines (BMs) are appealing candidates for powerful priors in variational autoencoders (VAEs), as they are capable of capturing nontrivial and multi-modal distributions over discrete variables. However, non-differentiability of the discrete units prohibits using the reparameterization trick, essential for lo…
New method for simplifying knots with specific properties.
We discuss replica analytic continuation using several simple models in order to prove mathematically the validity of replica analysis, which is used in a wide range of fields related to large scale complex systems. While replica analysis consists of two analytical techniques, the replica trick (or replica analytic con…
Paper develops a weighted linearization approach for vector fields.
Tricks adversarial attacks to target specific classes, improving classifier accuracy.
This paper solves a complex differential relation using a novel 'avoidance trick'.
A new network model combines features of DCBM, LSM, and β-model, using a cancellation trick for parameter estimation.
The reparameterization trick is widely used in variational inference as it yields more accurate estimates of the gradient of the variational objective than alternative approaches such as the score function method. Although there is overwhelming empirical evidence in the literature showing its success, there is relative…
Low-variance gradient estimation is crucial for learning directed graphical models parameterized by neural networks, where the reparameterization trick is widely used for those with continuous variables. While this technique gives low-variance gradient estimates, it has not been directly applicable to discrete variable…
Proves existence of planar curves with specific curvature.
The article confirms two quasi-alternating surgeries for 9 asymmetric L-space knots.
Improved diffusion bridge sampling with rKL-LD loss.
We stabilize the Kumaraswamy distribution for efficient sampling and differentiation.
The well-known Gumbel-Max trick for sampling from a categorical distribution can be extended to sample elements without replacement. We show how to implicitly apply this 'Gumbel-Top-' trick on a factorized distribution over sequences, allowing to draw exact samples without replacement using a Stochastic Beam Sea…