Percolation on complex networks has been used to study computer viruses, epidemics, and other casual processes. Here, we present conditions for the existence of a network specific, observation dependent, phase transition in the updated posterior of node states resulting from actively monitoring the network. Since tradi…
We define the information threshold in Bayesian decision-making.
problem Understanding the optimal amount of information for reliable classification.
method Defining the information threshold as the point of maximum curvature in the prior vs. posterior curve.
result At the information threshold, additional evidence does not significantly improve posterior probability.
The paper analyzes and proposes a new stopping criterion for recursive Bayesian classification.
problem Limitations of conventional stopping criteria in recursive Bayesian classification.
method Geometric interpretation of state posterior progression and analysis of conventional criteria.
result Proposes a new stopping criterion to overcome limitations of conventional methods.
High-dimensional VAEs inevitably collapse to prior, requiring large datasets for good performance.
problem Posterior collapse in VAEs leads to poor representation learning quality.
method Analyzed a minimal VAE in a high-dimensional limit, evaluating conditions for posterior collapse with respect to beta and dataset size.
result VAEs face 'inevitable posterior collapse' beyond a certain beta threshold, regardless of dataset size.
KL annealing helps VAEs avoid posterior collapse and overfitting.
problem Posterior collapse and overfitting in VAEs.
method Theoretical analysis of learning dynamics with KL annealing.
result Posterior collapse is inevitable when β exceeds a threshold. Paper proposes a method to improve variational inference for sparse networks.
problem Variational inference struggles with sparse networks, leading to inaccurate community detection.
method The method involves hard thresholding the posterior of community assignment after each iteration.
result The proposed method accurately recovers true community labels in sparse networks.
WBCP improves conformal prediction for distribution shifts using weighted Dirichlet posteriors.
problem Handling distribution shifts in conformal prediction.
method Generalizes Bayesian Quadrature Conformal Prediction (BQ-CP) to arbitrary importance-weighted settings.
result WBCP maintains coverage guarantees while providing richer uncertainty information.
We develop a novel approximate Bayesian computation (ABC) framework, ABCDP, that produces differentially private (DP) and approximate posterior samples. Our framework takes advantage of the Sparse Vector Technique (SVT), widely studied in the differential privacy literature. SVT incurs the privacy cost only when a cond…
The trade-off between the cost of acquiring and processing data, and uncertainty due to a lack of data is fundamental in machine learning. A basic instance of this trade-off is the problem of deciding when to make noisy and costly observations of a discrete-time Gaussian random walk, so as to minimise the posterior var…
This paper proposes a new algorithm for Gaussian process classification based on posterior linearisation (PL). In PL, a Gaussian approximation to the posterior density is obtained iteratively using the best possible linearisation of the conditional mean of the labels and accounting for the linearisation error. PL has s…
We propose an efficient meta-algorithm for Bayesian estimation problems that is based on low-degree polynomials, semidefinite programming, and tensor decomposition. The algorithm is inspired by recent lower bound constructions for sum-of-squares and related to the method of moments. Our focus is on sample complexity bo…
KSD Thinning uses KSD to thin MCMC samples efficiently.
problem Efficiently representing posterior distributions in Bayesian inference.
method KSD Thinning: retains only samples exceeding a KSD threshold.
result Established convergence and complexity tradeoffs for KSD Thinning.
Bayesian method estimates contamination factor for unsupervised anomaly detection.
problem No good methods for estimating contamination factor in unsupervised anomaly detection.
method Bayesian approach using mixture formulation of anomaly detector outputs.
result Estimated contamination factor distribution is well-calibrated and improves anomaly detection performance.
New method reduces GP bandit complexity while maintaining good performance.
problem Computational burden in Bayesian optimization with Gaussian processes.
method Information thresholding to compress GP posterior and reduce complexity.
result Sublinear regret bounds with sublinear posterior complexity.
Adversarial inference on tree models is possible with limited corruption, improving on Kesten-Stigum threshold.
problem Posterior inference on tree-structured graphical models in the presence of adversarial corruption.
method Dynamic programming via belief propagation, constrained adversarial corruption.
result Belief propagation can perform accurate inference with limited adversarial corruption.
We propose a vector-valued regression problem whose solution is equivalent to the reproducing kernel Hilbert space (RKHS) embedding of the Bayesian posterior distribution. This equivalence provides a new understanding of kernel Bayesian inference. Moreover, the optimization problem induces a new regularization for the …
New method improves posterior sampling for complex data models.
problem Sampling from posterior distributions in high-dimensional data.
method Tilted transport technique combining denoising oracle and log-likelihood.
result Boosted posterior is strongly log-concave, facilitating easier sampling.
In this paper we propose a model with a Dirichlet process mixture of gamma densities in the bulk part below threshold and a generalized Pareto density in the tail for extreme value estimation. The proposed model is simple and flexible allowing us posterior density estimation and posterior inference for high quantiles. …
Method selects the best deep learner for time-series prediction using Bayesian networks.
problem Selecting the most effective deep learning model for time-series prediction.
method Bayesian network selects deep learners based on input variables and cluster training data.
result Threshold value determines which deep learners predict time-series data robustly.
ABI bypasses likelihood intractability with nonparametric distribution matching.
problem Approximate Bayesian computation's inefficiency in high-dimensional settings and under diffuse priors.
method Adaptive Bayesian Inference (ABI) compares posterior distributions directly using nonparametric distribution matching and MSW distance.
result ABI significantly outperforms other methods in high-dimensional or dependent observation regimes.
Class imbalance presents a major hurdle in the application of data mining methods. A common practice to deal with it is to create ensembles of classifiers that learn from resampled balanced data. For example, bagged decision trees combined with random undersampling (RUS) or the synthetic minority oversampling technique…
Mixture models and topic models generate each observation from a single cluster, but standard variational posteriors for each observation assign positive probability to all possible clusters. This requires dense storage and runtime costs that scale with the total number of clusters, even though typically only a few clu…
For information retrieval and binary classification, we show that precision at the top (or precision at k) and recall at the top (or recall at k) are maximised by thresholding the posterior probability of the positive class. This finding is a consequence of a result on constrained minimisation of the cost-sensitive exp…
Conformal Bayes under label shift: post-hoc calibration vs. in-training adaptation
problem Bayesian prediction sets under label shift
method Post-hoc calibration vs. In-training adaptation
result Both strategies achieve valid coverage equally in an unbiased training regime
ESRL uses uncertainty quantification to learn safe, optimal policies in offline RL.
problem Challenges in interpreting and measuring uncertainty of learned policies in offline RL.
method Expert-Supervised Reinforcement Learning (ESRL) framework that uses hypothesis testing and posterior distributions.
result The framework can learn safe and optimal policies with theoretical guarantees and independent sample efficiency.
Two approaches improve conformal Bayes for label shift, one post-hoc and one in-training.
problem Improving prediction sets for target domain under label shift.
method Two complementary approaches: post-hoc calibration and in-training adaptation.
result In-training adaptation achieves up to 43% width reduction at unchanged coverage.
New acquisition functions improve Bernoulli LSE.
problem Efficiently estimating regions where a Bernoulli function is above or below a threshold.
method Developed new look-ahead acquisition functions for Gaussian process classification models.
result Demonstrated clear benefits of new acquisition functions on benchmark and real-world tasks.
We study the impact of learning on the optimal policy and the time-to-decision in an infinite-horizon Bayesian sequential decision model with two irreversible alternatives, exit and expansion. In our model, a firm undertakes a small-scale pilot project so as to learn, via Bayesian updating, about the project\textquoter…
We introduce thermodynamic response functions for singular Bayesian models.
problem Singular Bayesian models violate regular asymptotics due to non-identifiability and degenerate Fisher geometry.
method Posterior tempering induces thermodynamic response functions, linking WAIC, WBIC, and singular fluctuation.
result WAIC, WBIC, and singular fluctuation are unified within a thermodynamic response framework.
This paper treats prediction markets as Bayesian inverse problems to quantify uncertainty and identify event outcomes.
problem Uncertainty and identifiability in prediction market outcomes from price-volume histories.
method Formulates prediction markets as Bayesian inverse problems, introduces a log-odds observation model, and derives posterior uncertainty quantification and identifiability criteria.
result Explicit diagnostics for informative and stable inference regimes, and validation through synthetic data experiments.
We present here a general framework and a specific algorithm for predicting the destination, route, or more generally a pattern, of an ongoing journey, building on the recent work of [Y. Lassoued, J. Monteil, Y. Gu, G. Russo, R. Shorten, and M. Mevissen, "Hidden Markov model for route and destination prediction," in IE…
New local-search methods close the gap in sparse tensor PCA.
problem Sparse tensor PCA underperforms compared to other methods.
method Proposes new local-search methods including greedy and random-threshold variants.
result Proves local-search methods close the gap to best known polynomial-time procedures.
ABC method uses machine learning for likelihood-free inference.
problem Statistical inference in simulator-based models with intractable likelihoods.
method Direct comparison of empirical distributions via KL divergence estimator and contrastive learning.
result Asymptotic normality of ABC posterior distributions with properly scaled exponential kernel.
A keyword spotting (KWS) system determines the existence of, usually predefined, keyword in a continuous speech stream. This paper presents a query-by-example on-device KWS system which is user-specific. The proposed system consists of two main steps: query enrollment and testing. In query enrollment step, phonetic pos…
A novel ABC method for high-dimensional inverse problems using generative modeling and subset simulation.
problem Solving inverse-problems with high-dimensional inputs and expensive forward mappings.
method Joint deep generative modeling, Approximate Bayesian Computation (ABC) with Subset Simulation, and likelihood-free inference.
result Our method delivers promising performance without prior knowledge of the forward or noise distributions.
Bayesian approach for constructing and rebalancing sparse index-tracking portfolios.
problem Sparse tracking of a reference index with uncertainty quantification.
method Sparse linear regression with Laplace prior, empirical-Bayes calibration, Langevin-type MCMC, threshold-based rules.
result Posterior uncertainty on tracking error, portfolio composition, and rebalancing moves.
Stable training of deep normalizing flows for high-dimensional variational inference.
problem Training deep normalizing flows for high-dimensional posterior distributions is infeasible due to high stochastic gradient variance.
method Proposed a combination of soft-thresholding of scale and bijective soft log transformation to stabilize training.
result Stable training of Real NVPs for posterior distributions with thousands of dimensions is possible.
Researchers identify critical protein residues using advanced graph theory.
problem Identifying essential residues in proteins for function.
method Learning Random Geometric Graphs (RGG) with Cramer's V correlation and organic thresholding.
result Advanced RGG methods accurately identify critical residues compared to existing techniques.
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
Non-negative matrix factorization (NMF) is a knowledge discovery method that is used in many fields. Variational inference and Gibbs sampling methods for it are also wellknown. However, the variational approximation error has not been clarified yet, because NMF is not statistically regular and the prior distribution us…
Bayesian method discovers PDEs with variable coefficients robustly.
problem Discovering PDEs from noisy data is challenging.
method Bayesian sparse learning with tBGL-SS and Gibbs sampler.
result Method enhances robustness and model selection criteria.
We generalize the stochastic block model to the important case in which edges are annotated with weights drawn from an exponential family distribution. This generalization introduces several technical difficulties for model estimation, which we solve using a Bayesian approach. We introduce a variational algorithm that …
Bayesian models' singular fluctuation is shown to be akin to specific heat, influencing model complexity and generalization.
problem Understanding the thermodynamic interpretation of singular fluctuation in Bayesian models.
method Showed singular fluctuation as the curvature of Bayesian free energy and variance of log-likelihood observable under a Gibbs posterior.
result Singular fluctuation is the statistical analogue of specific heat, controlling model complexity and generalization.
Paper studies fair classification of functional data.
problem Mitigating disparities in functional data classification.
method Unified framework for fairness-aware functional classification.
result Established theoretical guarantees on fairness and excess risk controls.
An algorithmically hard phase was described in a range of inference problems: even if the signal can be reconstructed with a small error from an information theoretic point of view, known algorithms fail unless the noise-to-signal ratio is sufficiently small. This hard phase is typically understood as a metastable bran…
Bayesian nonparametric model predicts user activity and intervention success.
problem Predicting user activity and intervention success in online experiments.
method Bayesian nonparametric approach to model user heterogeneity and derive user activity predictions.
result The proposed method outperforms existing approaches in predicting user activity and intervention success.
Traditionally, community detection in graphs can be solved using spectral methods or posterior inference under probabilistic graphical models. Focusing on random graph families such as the stochastic block model, recent research has unified both approaches and identified both statistical and computational detection thr…
Study post-hoc Learning to Defer using density-ratio losses.
problem Optimizing decision-making between models and experts.
method Density-ratio losses for post-hoc L2D scorers, derived from class-probability estimation.
result The approach recovers known results and introduces new connections to expert comparison and anomaly detection.