The standard interpretation of importance-weighted autoencoders is that they maximize a tighter lower bound on the marginal likelihood than the standard evidence lower bound. We give an alternate interpretation of this procedure: that it optimizes the standard variational lower bound, but using a more complex distribut…
This paper compares gradient estimators in importance-weighted VI and justifies the superiority of DREP over REP.
problem Understanding the impact of gradient estimators on importance-weighted VI algorithms.
method Unified theoretical comparison of reparameterized and doubly-reparameterized gradient estimators tied to IWAE, VR, and VR-IWAE bounds.
result Formally justifies the superiority of doubly-reparameterized gradient estimators over reparameterized ones in importance-weighted VI.
New algorithm improves accuracy of importance weights for diverse applications.
problem Improving accuracy of importance weights for various applications.
method Formulated multicalibrated partitions and developed an efficient algorithm.
result Algorithm significantly improves accuracy of importance weights.
Improved discrete VAEs using relaxed Boltzmann priors for better performance.
problem Training discrete VAEs with tighter importance-weighted bounds.
method Two approaches for relaxing Boltzmann machines to continuous distributions, based on generalized overlapping transformations and the Gaussian integral trick.
result These relaxations outperform previous discrete VAEs with Boltzmann priors on MNIST and OMNIGLOT datasets.
Paper improves variance control in importance weighted variational bounds.
problem Improving the variance of gradient estimators for IWAE.
method Develops a novel control variate that grows SNR as √K for large K.
result Empirically, the method yields superior variance reduction for generative models.
Sharp analysis of out-of-distribution error in overparameterized models with importance weights.
problem Understanding and quantifying the degradation of performance in overparameterized models when faced with underrepresented data.
method Sharp analysis of an overparameterized Gaussian mixture model with spurious features and cost-sensitive interpolating solutions incorporating importance weights.
result Characterization of a novel tradeoff between worst-case robustness and average accuracy as a function of importance weight magnitude.
Paper formalizes and analyzes a new bound for variational inference.
problem Lack of theoretical guarantees in variational algorithms.
method Introduces VR-IWAE bound, a generalization of IWAE.
result VR-IWAE bound leads to unbiased gradient estimators.
Unified framework for analyzing pessimism in off-policy learning with regularized importance sampling.
problem High variance in importance weighting for off-policy learning.
method Unified PAC-Bayesian study of pessimism with regularized importance sampling.
result Derivation of a tractable PAC-Bayesian generalization bound for common importance weight regularizations.
Tighter ELBOs can harm learning, new algorithms improve performance.
problem Theoretical and empirical evidence shows that tighter ELBOs can reduce the signal-to-noise ratio of gradient estimators, hindering learning.
method Introduce three new algorithms: PIWAE, MIWAE, CIWAE, which improve over the standard IWAE.
result New algorithms can deliver improvements over IWAE, even when measured by IWAE's performance.
Hierarchical IWAE reduces sample redundancy for better inference.
problem Improving variational inference by reducing sample redundancy.
method Introduces a hierarchical structure to induce correlation among proposals.
result Maximizing the lower bound implicitly minimizes variance, improving inference performance.
Corrects distribution shift in target shift scenarios using importance weighting.
problem Analyzes importance weighting for correcting distribution shift under target shift.
method Analyzed importance-weighted kernel ridge regression under target shift.
result Shows that importance weighting corrects the train-test mismatch without altering input-space complexity.
New methods improve gradient estimation in autoencoders, enhancing generative network performance.
problem Improving gradient estimation in autoencoders to enhance learning.
method Developed and studied three methods: PIWAE, MIWAE, CIWAE.
result Generated approximate posterior distributions closer to true posterior distribution.
This work improves policy evaluation and selection using logarithmic smoothing for pessimistic off-policy estimation.
problem Offline evaluation and selection of policies from past data.
method Develops novel concentration bounds and a logarithmically smoothed estimator (LS) for improved policy selection and learning.
result The logarithmically smoothed estimator (LS) provides tighter bounds and better policy selection and learning.
Survey on importance weighting in machine learning applications.
problem Distribution shift in supervised learning.
method Weighting objective function or probability distribution based on instance importance.
result Importance weighting can guarantee desirable statistical properties in distribution shift scenarios.
A new method uses nearest neighbors for importance weighting.
problem Data covariate shift problems in machine learning.
method Nearest neighbor classification scheme for determining importance weights.
result Demonstrated effectiveness through comparative experiments on various classification tasks.
Improves probabilistic inference with new weighting method.
problem Improving variational inference for probabilistic models.
method Importance Weighted Variational Inference (IWVI) using augmented variational inference.
result IWVI is a practical technique for probabilistic inference.
The importance-weighted risk estimator can be skewed, leading to suboptimal regularization parameters.
problem Skewed sampling distribution of the importance-weighted risk estimator affects model selection.
method Empirical study of the sampling distribution of the importance-weighted risk estimator.
result The importance-weighted risk estimator produces overestimates for the majority of cases and underestimates for tail cases, leading to suboptimal regularization parameters.
A new method for evaluating and selecting policies in contextual bandits improves confidence intervals and policy quality.
problem Evaluating and selecting policies in contextual bandits with logged data.
method Self-normalized Importance Weighting (SN) estimator with Efron-Stein tail inequality and multiplicative bias control.
result The method provides tighter confidence intervals and better policy selection compared to competitors.
This study investigates the impact of importance weighting in deep learning models.
problem Understanding the effect of importance weighting in deep neural networks.
method The study uses theoretical and empirical approaches to analyze the behavior of importance weighting in deep learning models.
result Importance weighting impacts models early in training but diminishes over successive epochs in deep neural networks.
New loss function restores importance weighting in overparameterized models.
problem Restoring importance weighting in overparameterized neural networks.
method Introduced polynomially-tailed losses to restore effects of importance weighting.
result Polynomially-tailed losses improve performance in correcting distribution shift.
The variational autoencoder (VAE; Kingma, Welling (2014)) is a recently proposed generative model pairing a top-down generative network with a bottom-up recognition network which approximates posterior inference. It typically makes strong assumptions about posterior inference, for instance that the posterior distributi…
Paper improves REINFORCE for VI without restrictive assumptions.
problem Improves REINFORCE for VI without restrictive assumptions.
method Introduces VIMCO- ⋆ \star ⋆ gradient estimator to overcome SNR collapse. result VIMCO- ⋆ \star ⋆ achieves N \sqrt{N} N SNR scaling, superior to existing VIMCO. Improves cross-validation for biased data by adjusting risk estimator variance.
problem Cross-validation under sample selection bias produces suboptimal results.
method Introduces control variate to reduce variance of importance-weighted risk estimator.
result Control variate increases robustness to problematic weights.
Improved neural spike inference from calcium imaging data.
problem Neural spike inference from calcium imaging data.
method Importance weighted adversarial variational autoencoders (IWAE) with adversarial training.
result Adversarial IWAE methods outperform VAEs in inferring neural spikes.
A novel Bayesian computation method using importance weighting improves numerical stability and performance.
problem Bayesian computation stability and performance issues.
method Nonparametric approach via feature means, importance weighting, and kernel Bayes' rule.
result Importance weighted kernel Bayes' rule yields superior numerical stability and performance.
New method improves inference for hierarchical models.
problem Challenges in inference for large hierarchical models.
method Locally enhanced variational bounds with subsampling.
result Better posterior approximations than baselines.
We propose a sample-efficient alternative for importance weighting for situations where one only has sample access to the probability distribution that generates the observations. Our new method, called Geometric Resampling (GR), is described and analyzed in the context of online combinatorial optimization under semi-b…
We improve DGP models by using importance-weighted variational inference for better accuracy.
problem Accurate modeling of non-Gaussian marginals in deep Gaussian processes.
method Introduced noisy latent covariates and an importance-weighted objective for variational inference.
result The importance-weighted objective consistently outperforms classical variational inference, especially for deeper models.
We tackle causal inference under conditional moment restrictions using importance weighting.
problem Challenges in causal inference under conditional moment restrictions, especially in high-dimensional settings.
method Transform conditional moment restrictions to unconditional moment restrictions through importance weighting.
result Successfully estimate nonparametric functions defined under conditional moment restrictions.
Adapts score matching for missing data in flexible settings.
problem Learning data distribution with missing data.
method Adapted score matching to handle missing data, providing two approaches: importance weighting and variational.
result Variational approach performs best in high-dimensional settings.
U-statistics improve gradient estimation in importance-weighted variational inference.
problem High variance in gradient estimation for importance-weighted variational inference.
method Use U-statistics to average base gradient estimators on overlapping batches of size m, achieving lower variance.
result U-statistic variance reduction leads to modest to significant improvements in inference performance.
WR-CP reduces prediction set size and coverage gap under distribution shift.
problem Guaranteed coverage under distribution shift not achievable with i.i.d. assumption.
method Wasserstein distance, probability measure pushforwards, importance weighting, regularized representation learning.
result Reduces coverage gap to 3.2% across different confidence levels.
Paper tackles unbounded density ratio estimation for covariate shift adaptation.
problem Understudied challenge in statistical learning: unbounded density ratios.
method Three-step estimation method: relative density ratio, truncation, and transformation.
result Established rigorous convergence guarantees for density ratio and regression estimators.
Optimizes weights for better model performance in shifting data.
problem Improper importance weighting leads to poor model performance in data shifts.
method Interprets weights as a bias-variance trade-off and optimizes them simultaneously with model parameters.
result Optimizing weights significantly improves model generalization performance.
Improves model calibration and selection in unsupervised domain adaptation.
problem Distribution shifts in unsupervised domain adaptation.
method Developed a novel importance weighted group accuracy estimator.
result Improves state-of-the-art performances by 22% in model calibration and 14% in model selection.
A new one-step method for covariate shift adaptation.
problem Real-world data often violates the assumption of same distribution for training and test samples.
method Proposes a one-step optimization approach to jointly learn the model and weights.
result The proposed method achieves a generalization error bound and is empirically effective.
Corrects bias in learned generative models using likelihood-free importance weighting.
problem Bias in learned generative models relative to true data distribution.
method Estimate likelihood ratio using a classifier, apply importance weighting.
result Consistently improves goodness-of-fit metrics for deep generative models.
New method evaluates policies with latent confounders using optimal balance.
problem Evaluating policies with unobserved confounders in costly exploration scenarios.
method Importance weighting method to avoid latent outcome regression, minimizing adversarial balance objective.
result Provable consistency in policy evaluation with latent confounders, demonstrated empirically.
BR-SNIS reduces bias in self-normalized IS without increasing variance.
problem Bias in self-normalized IS.
method Iterated sampling-importance resampling (ISIR) to form a bias-reduced estimator.
result Significant reduction in bias without increasing variance.
The paper examines when importance weighting is needed for nonparametric and misspecified models.
problem When is importance weighting correction needed for covariate shift adaptation?
method Analysis of IW-corrected kernel ridge regression in various settings.
result The importance weighting correction is needed for nonparametric and misspecified models to obtain the best approximation of the true unknown function.
Proposes a method to transfer samples from source tasks to target tasks in RL.
problem Improving RL learning by selecting and weighting relevant samples from multiple tasks.
method Automatic estimation of importance weights for each source sample, applied to a batch RL algorithm.
result The proposed method achieves better learning performance and robustness to task differences.
A new method improves adversarial robustness by optimizing importance weights.
problem Adversarial training's non-uniform robustness across different data points.
method Doubly-robust instance reweighted adversarial training using distributionally robust optimization.
result Improves robustness against attacks on the weakest data points.
Generative models learn from biased data using weighted importance.
problem Learning from biased or related data distributions.
method Importance weighting to estimate loss with respect to target distribution.
result Effective in various settings with theoretical guarantees and good performance.
Improves transfer learning by weighting importance based on test-over-training density.
problem Distribution shift in training and test data.
method Joint and dynamic importance-predictor estimation, causal mechanism transfer.
result Enhanced transfer learning performance in complex, high-dimensional tasks.
New method optimizes model selection in high-dimensional regression models.
problem Model selection in high-dimensional misspecified regression models with covariate shift.
method Importance-weighted orthogonal greedy algorithm (IWOGA) and high-dimensional importance-weighted information criterion (HDIWIC).
result IWOGA + HDIWIC achieves optimal convergence rates in terms of prediction error.
IWeS selects examples by entropy-based importance sampling for subset selection.
problem Efficiently selecting examples for model training in batch settings.
method IWeS uses importance sampling based on model entropy to select examples.
result IWeS outperforms other subset selection algorithms on seven datasets.
The paper proves a new method to improve generalization in covariate-shift scenarios.
problem Improving performance on test distributions that differ from training distributions.
method Independence-driven importance weighting algorithms for feature selection.
result Theoretical proof that these algorithms can identify optimal variables for covariate-shift generalization.
Improves VAE training by refining variational parameters with BSVI.
problem Amortized inference in VAEs leads to suboptimal variational parameters and the amortization gap.
method Proposes BSVI, a refinement procedure using SVI's importance weights.
result Training VAEs with BSVI yields improved performance compared to SVI.