Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

219438656875 · Jun 202019922001200920182026
48 results for Wake-Sleep Algorithm

A new method for training generative models with sparse supervision.

problem Training deep generative models with sparse and varying supervision.
method Caffeinated Wake-Sleep (CWS) method, combining reweighted wake-sleep and teacher-forcing.
result The CWS method is robust to variable length supervision and performs well on various datasets.

New method learns flexible posterior distributions in generative models.

problem Learning accurate posterior distributions in hierarchical latent-variable models.
method Distributed distributional code Helmholtz machine with an extended wake-sleep algorithm.
result Outperforms state-of-the-art methods on various datasets.

DDVI uses diffusion models for variational inference, improving latent variable model performance.

problem Improving variational inference in latent variable models.
method Introduces diffusion-based variational posteriors trained with a regularized ELBO.
result Outperforms alternative variational posteriors on various benchmarks and a biology task.

A new approach to learning in brain-like networks using adversarial algorithms.

problem Complex inter-dependencies in brain-like networks not compatible with conditional independence assumptions.
method Adversarial algorithm for learning models of perceptual processing.
result The approach can mimic known neural phenomena and yields testable hypotheses.

Framework for lifelong learning using eigentasks to avoid forgetting and transfer knowledge.

problem Avoiding forgetting and transferring knowledge in lifelong learning.
method Eigentask framework: skills paired with generative models, wake-sleep cycle for learning and consolidation.
result Improved performance in supervised continual learning, evidence of forward knowledge transfer.

New method improves variational inference for better posterior approximation.

problem Challenges in minimizing inclusive KL divergence for amortized variational inference.
method Likelihood-tempered sequential Monte Carlo samplers to estimate inclusive KL gradient.
result SMC-Wake method fits variational distributions more accurately than existing methods.

QEM uses parallel importance weighting for fast approximate Bayesian inference.

problem Bayesian inference challenges in large models with many observations and latent variables.
method Expectation Maximization (EM) with massively parallel importance weighting.
result QEM is faster and more scalable than RWS and VI.

Highly expressive directed latent variable models, such as sigmoid belief networks, are difficult to train on large datasets because exact inference in them is intractable and none of the approximate inference methods that have been applied to them scale well. We propose a fast non-iterative approximate inference metho…

2014-01-31abs ↗pdf ↗

Unified framework for training generator, energy model, and inference model.

problem Training of generator, energy model, and inference model in a unified probabilistic formulation.
method Divergence Triangle framework integrating variational learning, adversarial learning, wake-sleep algorithm, and contrastive divergence.
result Unified training of generator, energy model, and inference model without costly Markov chain Monte Carlo methods.

InfoGAN learns disentangled representations without supervision.

problem Learning interpretable representations without labeled data.
method Generative adversarial network with mutual information maximization.
result InfoGAN successfully disentangles various latent variables from observations.

AISLE framework improves on IWAE by directly optimising proposal distribution.

problem IWAE's multi-sample objective leads to inference-network gradients that break down with increasing samples.
method Introduces AISLE framework, which optimises proposal distribution directly.
result AISLE admits IWAE-STL and IWAE-DREG as special cases, avoiding breakdown.

Paper improves variance control in importance weighted variational bounds.

problem Improving the variance of gradient estimators for IWAE.
method Develops a novel control variate that grows SNR as √K for large K.
result Empirically, the method yields superior variance reduction for generative models.

Develops a new method for efficient probabilistic inference.

problem Efficient inference for models with dynamic computation graphs.
method Introduces combinator library for Probabilistic Torch framework.
result Models can be trained using stochastic methods that optimize variational or wake-sleep objectives.

Examines algorithmic modeling across three cultures.

problem Tackles algorithmic modeling in different cultural contexts.
method Uses parametric regressions, interpretable algorithms, and complex algorithms.
result Extension of Leo Breiman's thesis to include cultural differences.

Meta-algorithm selection aims to choose the best algorithm selector for a given problem instance.

problem Selecting the best algorithm selector for a specific problem instance.
method Apply algorithm selection to the selection of other algorithms (meta-algorithm selection).
result Meta-algorithm selection can be beneficial in some cases but faces challenges in solving the meta-level problem.

Combines multiple bandit algorithms to create a nearly optimal single algorithm.

problem Designing a single bandit algorithm that performs nearly as well as the best individual algorithm in a stochastic environment.
method Develops two general corralling algorithms that achieve favorable regret guarantees.
result The regret of the corralling algorithms is no worse than the best individual algorithm's performance.

New algorithms improve stochastic optimization and online learning efficiency.

problem Efficient optimization and online learning algorithms for stochastic problems.
method Accelerated randomized coordinate descent algorithms.
result Significantly less per-iteration complexity and better regret performance.

The exchange algorithm is studied for its convergence and asymptotic variance.

problem Theoretical limitations of the exchange algorithm in sampling from doubly-intractable distributions.
method Theoretical analysis of the exchange algorithm's convergence speed and asymptotic variance.
result The exchange algorithm converges at a geometric rate and satisfies a Central Limit Theorem.

New algorithms optimize algorithm parameters in online settings with reduced computational costs.

problem Optimizing algorithm parameters in online settings with volatile and discontinuous losses.
method Developed semi-bandit optimization algorithms that leverage extra information to reduce computational costs.
result Achieved regret bounds as good as full-information feedback with significantly less computational effort.

Improves algorithm selection for thousands of candidates using dyadic features.

problem Selecting the best algorithm from a large set of candidates for specific problems.
method Proposes extreme algorithm selection (XAS) with dyadic feature representation.
result Improves significantly over current state of the art in various metrics.

AIDE measures the accuracy of probabilistic inference algorithms.

problem Measuring the accuracy of approximate inference algorithms on specific data sets.
method AIDE is an algorithm based on viewing inference algorithms as probabilistic models and auxiliary variables.
result AIDE captures the qualitative behavior of inference algorithms and detects failure modes.

New algorithm improves worst Value-at-Risk computation for risky portfolios.

problem Computing worst Value-at-Risk in heterogeneous portfolios is numerically challenging.
method Introduced an Adaptive Rearrangement Algorithm to improve the Rearrangement Algorithm.
result The Adaptive Rearrangement Algorithm provides more accurate approximations of worst Value-at-Risk.

New algorithms decode Markov chains with near-optimal performance, even with small latency.

problem Online decoding of nthn^{th} order ergodic Markov chains with latency constraints.
method Deterministic and randomized algorithms using dynamic programs, with lower bounds established.
result Near-optimal performance of algorithms with minimal latency, outperforming existing methods.

New ELM algorithms reduce computation time and complexity.

problem Efficient computation of extreme learning machine (ELM) algorithms.
method Developed inverse-free ELM algorithms using recursive matrix inverse and inverse LDL' factorization.
result Proposed algorithms significantly reduce computational complexity.