Faster and versatile sampler for Bayesian linear regression.
problem Efficiently sampling from Bayesian linear regression models with arbitrary priors.
method Slice sampler exploiting linear regression likelihood structure.
result Better effective sample size per second than alternatives.
Nested sampling improved for arbitrary priors.
problem Technical obstacle to using nested sampling with arbitrary priors.
method Parametric bijectors trained on samples from a desired prior density.
result Nested sampling can be used with arbitrary priors.
New approach to neural networks by incorporating observation noise and arbitrary prior means.
problem Misspecification on noisy data and limitations of NTK-GP equivalence.
method Introducing a regularizer for observation noise and proposing a shifted network for arbitrary prior means.
result Removes key obstacles to practical Gaussian process modeling in neural networks.
Unified algorithm for incorporating various prior knowledge in multiple testing.
problem Improving power and precision in multiple testing procedures with prior knowledge.
method p-filter algorithm that incorporates four types of prior knowledge: null hypotheses, penalties, groups, and independence.
result Unified framework allows for recovery of various known algorithms.
Algorithm learns from both labeled and arbitrary test examples, giving guarantees for bounded VC dimension classes.
problem Learning from arbitrary test examples, not just perturbations.
method Selective transductive learning algorithm that outputs abstaining predictions.
result Nontrivial guarantees for bounded VC dimension classes with arbitrary train and test distributions.
Proposes a new technique for VAEs using manifold-valued variables.
problem Effect of prior distribution choice on VAE learning capacity.
method Embedding-reparameterization procedure (ER) for manifold-valued latent variables.
result ER technique outperforms conventional VAE on a toy benchmark.
While Bayesian methods are praised for their ability to incorporate useful prior knowledge, in practice, convenient priors that allow for computationally cheap or tractable inference are commonly used. In this paper, we investigate the following question: for a given model, is it possible to compute an inference result…
The study finds that memorization is necessary or harmful depending on the prior distribution and noise level.
problem The impact of memorization on generalization in overparameterized models.
method An overparameterized linear model with general priors in a Bayesian setup.
result Explicit conditions for optimal generalization based on the prior distribution and noise level.
New static vacuum metrics confirmed for near Euclidean boundary data.
problem Establishing sufficient conditions for near Euclidean boundary data in static vacuum metrics.
method Using new arguments from studying the conjecture for arbitrary static vacuum metrics.
result Any hypersurface in a dense subfamily is static regular.
Study improves convergence rates for GVI under prior misspecification.
problem Improving convergence rates for GVI under prior misspecification.
method Proves rates of convergence and robustness to prior misspecification in GVI framework.
result Establishes sufficient conditions for existence and uniqueness of GVI posteriors.
Sharp asymptotics derived for phase retrieval and compressed sensing with random generative priors.
problem Phase retrieval and compressed sensing with random measurement matrices.
method Sharp asymptotics derived for optimal performance and polynomial algorithm for random generative priors.
result Compressed phase retrieval becomes tractable with random generative priors, unlike sparse priors.
WS diffusion models handle anisotropic Gaussian noise better than conventional methods.
problem Handling anisotropic Gaussian noise in imaging inverse problems.
method Whitened Score (WS) diffusion models based on stochastic differential equations.
result WS DMs outperform conventional DMs on anisotropic Gaussian noise.
New method learns SIMs with arbitrary monotone activations without strong distributional assumptions.
problem Learning Single-Index Models with arbitrary monotone activations.
method Based on omniprediction with calibrated multiaccuracy and Bregman divergences.
result First agnostic learning result for SIMs with arbitrary monotone activations.
K-priors enable quick adaptation with minimal retraining.
problem Machine learning models struggle to adapt to changes efficiently.
method Combines weight and function-space priors to reconstruct past gradients.
result Adaptation with K-priors achieves similar performance to full retraining with less data.
Exemplar-based clustering methods have been shown to produce state-of-the-art results on a number of synthetic and real-world clustering problems. They are appealing because they offer computational benefits over latent-mean models and can handle arbitrary pairwise similarity measures between data points. However, when…
We consider the problem of learning from distributed data in the agnostic setting, i.e., in the presence of arbitrary forms of noise. Our main contribution is a general distributed boosting-based procedure for learning an arbitrary concept space, that is simultaneously noise tolerant, communication efficient, and compu…
The paper extends and applies a new shrinkage prior in Bayesian factor analysis.
problem Estimating the number of factors in sparse Bayesian factor analysis.
method Introduces and extends a generalized cumulative shrinkage process (CUSP) prior.
result Exchangeable spike-and-slab shrinkage priors imply increasing shrinkage as the column index increases.
New method improves Robbins-Monro algorithm convergence with prior information.
problem Improving convergence speed of Robbins-Monro algorithm.
method Integrates prior information into Robbins-Monro iteration without regression model.
result Prior-information Robbins-Monro sequence converges faster than standard.
Study proposes learning optimal priors from data for better Bayesian inference.
problem Challenges the use of noninformative uniform priors in Bayesian inference.
method Machine learning approach to learn optimal priors from data using a target function.
result Study models consistently outperformed baseline models in Wikipedia category classification.
Algorithm learns RBMs with arbitrary external fields, improving on previous constraints.
problem Learning RBMs with arbitrary external fields, improving on previous constraints.
method Greedy algorithm that maximizes covariance between observed nodes sharing latent neighbors.
result Algorithm can learn RBMs with arbitrary external fields, improving on previous constraints.
New algorithm minimizes cumulative loss in dynamic linear bandits without prior knowledge of comparator switches.
problem Minimizing cumulative loss in dynamic linear bandits with unknown number of switches.
method Combining several bandit algorithms to adapt to unknown number of switches without prior knowledge.
result First algorithm achieving optimal regret guarantee of O ( d ( 1 + S T ) T ) \mathcal{O}\big(\sqrt{d(1+S_T) T}\big) O ( d ( 1 + S T ) T ) up to poly-logarithmic terms. Bayesian neural networks' performance varies with prior choice, affecting their ability to identify unknowns.
problem The impact of prior choice on Bayesian neural networks' ability to identify unknowns.
method Evaluation of different prior distributions on classification tasks using BNNs and NNs with Monte Carlo dropout.
result Prior choice significantly impacts BNNs' ability to identify unknowns, affecting true and false positive rates.
Extends variable screening for ultrahigh-dimensional models, reducing dimensionality to sample size.
problem Statistical inference challenges in ultrahigh-dimensional linear models.
method Extends correlation-based variable screening to arbitrary linear models and post-screening inference techniques.
result Shows a condition (screening condition) sufficient for successful variable screening in arbitrary linear models.
Recent advances in Bayesian reinforcement learning (BRL) have shown that Bayes-optimality is theoretically achievable by modeling the environment's latent dynamics using Flat-Dirichlet-Multinomial (FDM) prior. In self-interested multi-agent environments, the transition dynamics are mainly controlled by the other agent'…
LDTA expands LDA's topic modeling capacity with tree-structured priors.
problem Limited expressiveness of Dirichlet priors in LDA for complex topic relationships.
method Introduces Latent Dirichlet-Tree Allocation (LDTA) with Dirichlet-Tree (DT) priors, and develops universal mean-field variational inference and Expectation Propagation.
result LDTA enables expressive, tree-structured priors over topic proportions, expanding modeling capacity of LDA.
Paper tackles training binary classifiers from unlabeled data with minimal supervision.
problem Training arbitrary binary classifiers from only unlabeled data is impossible without supervision.
method Proposes an ERM-based learning method from two sets of unlabeled data with different class priors.
result The proposed method is consistent and outperforms state-of-the-art methods.
Accelerates Bayesian optimization using learned weight-prior.
problem Optimizing expensive functions with limited auxiliary data.
method Constructs a GP covariance from auxiliary data to model a more appropriate weight prior.
result Accelerates Bayesian optimization on test functions and real-world applications.
PGD algorithms solve nonlinear inverse problems with generative priors using noisy measurements.
problem Signal estimation from noisy nonlinear measurements with generative priors.
method Projected gradient descent algorithms for two cases: unknown and known nonlinearity.
result PGD algorithms converge linearly to optimal statistical rates using arbitrary initialization.
TzK model learns tight conditional priors using side information.
problem Learning tight conditional priors using side information.
method TzK is a conditional probability flow-based model that exploits attributes to learn tight conditional priors around target observations. It trains via approximated ML and supports supervised, unsupervised, and semi-supervised learning.
result TzK produces efficient and stable approximations of arbitrary data distributions, comparable to state-of-the-art models.
Unified analysis of privacy leakage in correlated data considering prior knowledge.
problem Understanding the impact of prior knowledge and data correlation on privacy leakage.
method Proposed prior differential privacy (PDP) and analyzed using WHG and multivariate Gaussian models.
result Derived closed-form expression for continuous data and chain rule for discrete data.
New examples show some convex-cocompact subgroups are separable.
problem Whether all convex-cocompact subgroups are separable.
method Using Manning-Mj-Sageev construction, examples of separable subgroups of arbitrary finite rank are given.
result Examples of separable convex-cocompact subgroups of arbitrary finite rank exist.
Bayesian evidence computation revisited for model selection with improper priors.
problem Model selection with improper priors and their impact on Bayesian evidence computation.
method Employing improper priors in model selection problems, distinguishing between Bayesian evidence and fake evidences.
result Diffuse priors asymptotically to infinity do not recover the area under the likelihood.
Fast Bayesian inference with adaptable priors for real-time applications.
problem Intractable exact posterior computation limits Bayesian inference's adoption.
method Distribution Transformer architecture that learns mappings between priors and posteriors.
result Significant reduction in computation time from minutes to milliseconds.
Let U U U be an arbitrary word in letters x 1 ± 1 , . . . , x m ± 1 x_1^{\pm 1}, ..., x_m^{\pm 1} x 1 ± 1 , ... , x m ± 1 and m ≥ 2 m \ge 2 m ≥ 2 . We prove that the group presentation < x 1 , . . . , x m ∥ U x i U − 1 = x i + 1 , i = 1 , . . . , m − 1 > <x_1, ..., x_m \|\ U x_i U^{-1} = x_{i+1}, i=1,..., m-1> < x 1 , ... , x m ∥ U x i U − 1 = x i + 1 , i = 1 , ... , m − 1 > is aspherical. The proof is based upon prior partial results of A. Klyachko and the author on the asphericity of such presentations.
We solve image inverse problems using a flow-based noise model.
problem Image inverse problems with complex noise patterns.
method Normalizing flow prior for maximum a posteriori estimation.
result Empirical validation on various inverse problems.
Neural-g models mixtures of densities with flexibility and accuracy.
problem Accurately estimating prior densities in g g g -models. method Neural network with softmax output for valid probability densities.
result Neural-g captures various prior shapes including flat, heavy-tailed, and discontinuous.
The study examines the retrieval capabilities of RBMs and generalized Hopfield networks under various prior distributions.
problem Characterizing the state of RBMs and Hopfield networks under different prior distributions.
method Equivalence between RBMs and generalized Hopfield networks, analysis of phase transitions, and study of retrieval capabilities.
result The retrieval phase is robust and exists at low load for every pattern distribution.
Transformer pretraining yields strong EB performance without explicit adaptation.
problem Empirical Bayes problems with unknown test distributions.
method Indirect analysis of pretrained transformer's performance under universal priors.
result Near-optimal regret bound of O ~ ( 1 n ) \widetilde{O}(\frac{1}{n}) O ( n 1 ) for arbitrary test distributions. These notes form part of a lecture course on gauge theory. The material covered is standard in the physics literature, but perhaps less well-known to mathematicians. The purpose of these notes is to make spontaneous symmetry breaking and the Higgs mechanism of mass generation for elementary particles more easily access…
Proposes a method to integrate prior knowledge into trajectory prediction models.
problem Improving accuracy and robustness in trajectory prediction models.
method Continual learning approach that allows integration of arbitrary prior knowledge and probabilistic predictions.
result Outperforms non-informed and informed learning methods, using half as many observation examples.
Two new estimators improve VAE training for hierarchical and prior parameters.
problem Efficient gradient estimation for VAEs with hierarchical and prior parameters.
method Developed two generalizations of Doubly-Reparameterized Gradient Estimators (DReGs) for VAEs.
result Improved training of conditional and hierarchical VAEs on image modeling tasks.
New bounds for SA with arbitrary norm contractions and Markovian noise.
problem Finite-time analysis of two-time-scale stochastic approximation with arbitrary norm contractions and Markovian noise.
method Use of generalized Moreau envelope for arbitrary norm contractions and solutions of Poisson equation for Markovian noise.
result Mean square error decays at rates of O ( 1 / n 2 / 3 ) O(1/n^{2/3}) O ( 1/ n 2/3 ) and O ( 1 / n ) O(1/n) O ( 1/ n ) under different conditions. We integrate camera pose correlations into deep models using Gaussian processes.
problem Lack of inter-frame reasoning in deep neural networks.
method Derive a principled framework combining camera pose information with deep models using a novel view kernel.
result Soft-prior knowledge aids pose-related vision tasks like novel view synthesis.
Proves a new law of robustness for interpolating arbitrary data distributions.
problem Understanding robust interpolation for arbitrary data distributions.
method Proves a Lipschitzness lower bound for robust interpolation.
result Demonstrates a two-fold law of robustness for interpolating functions.
e-LOND algorithm controls FDR in online testing with arbitrary dependencies.
problem Online testing of hypotheses with unknown dependencies.
method e-LOND algorithm for FDR control under arbitrary dependence.
result e-LOND provides more power than existing methods through simulations.
KITT uses transformers to quickly recommend kernels for GP models.
problem Kernel selection for high-dimensional GP regression models.
method Transformer-based architecture for generating kernel recommendations.
result KITT selects kernels that perform well on various regression benchmarks.
Study optimal asset liquidation under uncertain drift and volatility changes.
problem Optimal liquidation of assets with unknown drift and stochastic volatility.
method Modelled as a four-dimensional optimal stopping problem, solved using filtering theory and approximating sequences of three-dimensional problems.
result Determined optimal liquidation strategy and structural properties.
In this paper we address the problem of learning the structure of a Bayesian network in domains with continuous variables. This task requires a procedure for comparing different candidate structures. In the Bayesian framework, this is done by evaluating the {em marginal likelihood/} of the data given a candidate struct…