AirRL uses RL to infer urban air quality from selected stations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DGPs improve air quality inference from sparse data.
Federated learning leaks participant dataset quality even with secure aggregation.
Inferring air quality from a limited number of observations is an essential task for monitoring and controlling air pollution. Existing inference methods typically use low spatial resolution data collected by fixed monitoring stations and infer the concentration of air pollutants using additional types of data, e.g., m…
Bayesian Neural Networks (BNNs) place priors over the parameters in a neural network. Inference in BNNs, however, is difficult; all inference methods for BNNs are approximate. In this work, we empirically compare the quality of predictive uncertainty estimates for 10 common inference methods on both regression and clas…
Enhances statistical inference using synthetic data.
SDM Policy accelerates inference for robotic tasks while maintaining high action quality.
FAB-PPI uses prior knowledge to improve prediction-powered inference.
Statistical inference on graphs is a burgeoning field in the applied and theoretical statistics communities, as well as throughout the wider world of science, engineering, business, etc. In many applications, we are faced with the reality of errorfully observed graphs. That is, the existence of an edge between two vert…
FPPI selectively uses predictions to improve inference efficiency.
We introduce a novel generative autoencoder network model that learns to encode and reconstruct images with high quality and resolution, and supports smooth random sampling from the latent space of the encoder. Generative adversarial networks (GANs) are known for their ability to simulate random high-quality images, bu…
Adaptive workflow combines fast amortized inference with MCMC for many datasets.
TRAiL is a linear bandit algorithm that ensures optimal regret and guarantees inference quality.
PGMax automates PGM inference on GPUs, improving quality and speed.
Generative model improves sample quality on ImageNet32.
Inference is an integral part of probabilistic topic models, but is often non-trivial to derive an efficient algorithm for a specific model. It is even much more challenging when we want to find a fast inference algorithm which always yields sparse latent representations of documents. In this article, we introduce a si…
New method improves inference for discrete diffusion models, achieving better quality and efficiency.
This paper analyzes speculative decoding, a method to speed up large language model inferences.
URGE improves diffusion model quality without gradients or Hessian.
The pairwise influence matrix of Dobrushin has long been used as an analytical tool to bound the rate of convergence of Gibbs sampling. In this work, we use Dobrushin influence as the basis of a practical tool to certify and efficiently improve the quality of a discrete Gibbs sampler. Our Dobrushin-optimized Gibbs samp…
One of the core problems in statistical models is the estimation of a posterior distribution. For topic models, the problem of posterior inference for individual texts is particularly important, especially when dealing with data streams, but is often intractable in the worst case. As a consequence, existing methods for…
Remasking improves the quality of discrete diffusion models for natural language and image generation.
DriftLite improves inference quality of diffusion models without retraining.
CAFL breaks feedback loops in recommender systems using causal inference.
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
New algorithm improves inference-time alignment without reward hacking.
This work explores function-space inference using KL divergence and proposes Bayesian linear regression as a benchmark.
Nested Chinese Restaurant Process (nCRP) topic models are powerful nonparametric Bayesian methods to extract a topic hierarchy from a given text corpus, where the hierarchical structure is automatically determined by the data. Hierarchical Latent Dirichlet Allocation (hLDA) is a popular instance of nCRP topic models. H…
Gaussian processes (GPs) offer a flexible class of priors for nonparametric Bayesian regression, but popular GP posterior inference methods are typically prohibitively slow or lack desirable finite-data guarantees on quality. We develop an approach to scalable approximate GP regression with finite-data guarantees on th…
IMM generates high-quality samples in few steps with stable training.
We consider causal inference in the presence of unobserved confounding. We study the case where a proxy is available for the unobserved confounding in the form of a network connecting the units. For example, the link structure of a social network carries information about its members. We show how to effectively use the…
A key advance in learning generative models is the use of amortized inference distributions that are jointly trained with the models. We find that existing training objectives for variational autoencoders can lead to inaccurate amortized inference distributions and, in some cases, improving the objective provably degra…
The automation of posterior inference in Bayesian data analysis has enabled experts and nonexperts alike to use more sophisticated models, engage in faster exploratory modeling and analysis, and ensure experimental reproducibility. However, standard automated posterior inference algorithms are not tractable at the scal…
Gibbs sampling is a workhorse for Bayesian inference but has several limitations when used for parameter estimation, and is often much slower than non-sampling inference methods. SAME (State Augmentation for Marginal Estimation) \cite{Doucet99,Doucet02} is an approach to MAP parameter estimation which gives improved pa…
Fast approximate inference for non-Gaussian data.
Approximate Bayesian Computation is widely used in systems biology for inferring parameters in stochastic gene regulatory network models. Its performance hinges critically on the ability to summarize high-dimensional system responses such as time series into a few informative, low-dimensional summary statistics. The qu…
Framework evaluates quality of synthetic data generated with differential privacy.
IterefinE combines KG refinement with embeddings to improve KG quality.
Privacy-preserving inference for clinical trials using differential privacy.
This paper compares self-reflection and budget tuning for LLMs, revealing domain-specific performance gains.
New model clusters mixed-type data with missing values, improving air quality analysis.
A classical problem in causal inference is that of matching, where treatment units need to be matched to control units based on covariate information. In this work, we propose a method that computes high quality almost-exact matches for high-dimensional categorical datasets. This method, called FLAME (Fast Large-scale …
This work improves neural likelihood surrogates for stochastic models with a score-augmented loss.
Improves model classification accuracy in black-box settings.
Bayesian approach improves image captioning quality metrics.
Stochastic variational inference allows for fast posterior inference in complex Bayesian models. However, the algorithm is prone to local optima which can make the quality of the posterior approximation sensitive to the choice of hyperparameters and initialization. We address this problem by replacing the natural gradi…
New method simplifies Bayesian analysis for categorical data.
The mean field methods, which entail approximating intractable probability distributions variationally with distributions from a tractable family, enjoy high efficiency, guaranteed convergence, and provide lower bounds on the true likelihood. But due to requirement for model-specific derivation of the optimization equa…