Bayesian inference for deep neural networks using trace-class priors and MLMC.
problem Efficient Bayesian inference for deep neural networks.
method Trace-class neural network priors and Multilevel Monte Carlo method.
result Optimal computational complexity for Bayesian inference of TNN models.
New BNN architectures reduce computational cost for uncertainty quantification.
problem High computational cost in Bayesian neural networks.
method Partial trace-class Bayesian neural networks (PaTraC BNNs).
result Comparable uncertainty quantification with fewer parameters.
Develops trace class operators and inverse Laplacian theory for infinite dimensions.
problem Understanding trace class operators and inverse Laplacian on infinite dimensional spaces.
method Presentation of trace class operators and construction of inverse Laplacian on closed manifolds.
result Original trace computations involving the inverse Laplacian on the torus.
This work presents a parametrized family of divergences, namely Alpha-Beta Log- Determinant (Log-Det) divergences, between positive definite unitized trace class operators on a Hilbert space. This is a generalization of the Alpha-Beta Log-Determinant divergences between symmetric, positive definite matrices to the infi…
New Gaussian priors for neural networks improve scalability and Bayesian inference stability.
problem Scalability and stability issues in Bayesian neural network inference.
method Introduces a new Gaussian neural network prior with decreasing variance in network width, enabling stable MCMC sampling.
result The new prior enables stable MCMC sampling for Bayesian neural network inference, improving scalability and stability.
Maps with many singularities found in complex space.
problem Constructing maps with singularities in complex space.
method Created a map with infinitely many Schoen-Wolfson singularities on a disc.
result Found a Ck map with smooth trace in C2. We study families of Dirac-type operators, with compatible perturbations, associated to wedge metrics on stratified spaces. We define a closed domain and, under an assumption of invertible boundary families, prove that the operators are self-adjoint and Fredholm with compact resolvents and trace-class heat kernels. We …
A cocycle Ω:P×G→H taking values in a Lie group H for a free right action of G on P defines a principal bundle Q with the structure group H over P/G. The Chern character of a vector bundle associated to Q defines then characteristic classes on X. This observation becomes useful in the case …
We consider isotropic Lévy processes on a compact Riemannian manifold, obtained from an Rd-valued Lévy process through rolling without slipping. We prove that the Feller semigroups associated with these processes extend to strongly continuous contraction semigroups on Lp, for 1≤p<∞, and that t…
Infinite-dimensional SBDMs improve image generation across multiple resolutions.
problem Efficient image generation at high resolutions and across different levels.
method Developed SBDMs in infinite-dimensional setting, using trace class operators and operator networks.
result Improved efficiency and generalization across resolution levels.
Let (X,h) be a compact and irreducible Hermitian complex space of complex dimension v>1. In this paper we show that the Friedrichs extension of both the Laplace-Beltrami operator and the Hodge-Kodaira Laplacian acting on functions has discrete spectrum. Moreover we provide some estimates for the growth of the corre…
Abstract: Determinants and formulas for operators on various spaces.
problem Determinants and formulas for operators on different algebras and spaces.
method Use of Poincaré type determinants, invariant operators, and full matrix-symbols.
result Explicit formulas for determinants of elliptic operators and periodic pseudo-differential operators.
Let A0 and A1 be two self-adjoint Fredholm Dirac-type operators defined on two non-compact manifolds. If they coincide at infinity so that the relative heat operator is trace-class, one can define their relative eta function as in the compact case. The regular value of this function at the zer…
We construct a Hennings type logarithmic invariant for restricted quantum sl(2) at a 2p-th root of unity. This quantum group U is not braided, but factorizable. The invariant is defined for a pair: a 3-manifold M and a colored link L inside M. The link L is split into two parts colored…
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
problem Computing divergence between Gaussian measures in infinite-dimensional Hilbert space.
method Closed form expression and regularization for divergence calculation.
result Closed form expression and regularization for Geometric Jensen-Shannon divergence.
A new model captures forward curve dynamics with stochastic volatility.
problem Modeling continuous-time evolution of forward curves in financial markets.
method Affine stochastic volatility model with modulated dynamics.
result Model allows for maturity-specific risk and volatility clustering.
This paper solves nonparametric estimation of continuous DPPs using kernel methods.
problem Estimating continuous Determinantal Point Processes (DPPs) without assuming a parametric form.
method Developed a fixed point algorithm based on a representer theorem for nonnegative functions in RKHS.
result Demonstrated a finite-dimensional problem for nonparametric MLE of continuous DPPs.
Let V⊂CPn be an irreducible complex projective variety of complex dimension v and let g be the Kähler metric on $\reg(V)$, the regular part of V, induced by the Fubini Study metric of CPn. In this setting Li and Tian proved that $W^{1,2}_0(\reg(V),g)=W^{1,2}(\reg(V…
The paper introduces Causal Neural Operators to approximate operators in stochastic analysis.
problem Leveraging temporal structure in non-linear operators for deep learning models.
method Designing a deep learning model framework for infinite-dimensional linear metric spaces.
result Causal Neural Operators can uniformly approximate Hölder or smooth trace class operators.
The paper advances U-statistics in dependent settings, improving spectral estimation and goodness-of-fit tests.
problem Non-asymptotic analysis of U-statistics in dependent Markov chain settings.
method Proved new concentration and exponential inequalities for U-statistics, applied to spectral estimation, online algorithms, and goodness-of-fit tests.
result Established new results for spectral estimation, online algorithms, and goodness-of-fit tests in Markov chain settings.
Develops hypothesis tests for conditional distributions using learning-theoretic bounds.
problem Testing differences in conditional distributions and functionals.
method Transforming learning-theoretic bounds into hypothesis tests for conditional expectations.
result Establishes comprehensive foundation for conditional testing, including theoretical guarantees and practical implementations.
Let (X,h) be a compact and irreducible Hermitian complex space of complex dimension m. In this paper we are interested in the Dolbeault operator acting on the space of L2 sections of the canonical bundle of reg(X), the regular part of X. More precisely let $\overline{\mathfrak{d}}_{m,0}:L^2Ω^{m,0}(reg(X),h)\…
Study convergence and approximations of entropic regularized Wasserstein distances for Gaussian and RKHS measures.
problem Convergence and approximations of entropic regularized Wasserstein distances in Gaussian and RKHS settings.
method Analysis of convergence and finite sample approximations of entropic regularized Wasserstein distances in Gaussian and RKHS settings.
result Strictly weaker convergence in 2-Sinkhorn divergence for Gaussian measures compared to exact 2-Wasserstein distance.
A new Weyl prior is proposed for Bayesian statistics, offering a more canonical choice for parameter α.
problem Choosing a prior distribution for Bayesian inference.
method Proposed a new Weyl prior based on the Weyl structure on a statistical manifold.
result The Weyl prior is a special case of the α-parallel prior with α = -n, where n is the dimension of the statistical manifold.
Improves AI-prior reliability for Bayesian inference.
problem Error propagation from predictive models into posterior inference.
method Rectified AI-informed prior elicitation framework.
result Significant reduction in bias and improvement in predictive performance.
Informative Bayesian priors are often difficult to elicit, and when this is the case, modelers usually turn to noninformative or objective priors. However, objective priors such as the Jeffreys and reference priors are not tractable to derive for many models of interest. We address this issue by proposing techniques fo…
While Bayesian methods are praised for their ability to incorporate useful prior knowledge, in practice, convenient priors that allow for computationally cheap or tractable inference are commonly used. In this paper, we investigate the following question: for a given model, is it possible to compute an inference result…
Researchers derive exact priors for finite Bayesian neural networks.
problem Understanding non-Gaussian priors in finite Bayesian neural networks.
method Analytical derivation of function space priors for finite fully-connected feedforward networks.
result Exact solutions for priors of finite networks, including Meijer G-function for linear networks and mixtures for ReLU networks.
PRCD-MAP learns to trust imperfect priors in causal discovery, improving accuracy and robustness.
problem Tackles the brittle trade-off between blind trust and rejection of external priors in causal discovery.
method Proposes PRCD-MAP, a soft prior-consumption layer that assigns per-edge trust to imperfect priors and modulates regularization in a MAP objective.
result Enjoys a population-level safety guarantee and outperforms existing methods on real-world causal discovery tasks.
Bayesian metalearning improves performance in linear bandits with misspecified priors.
problem Improper priors lead to suboptimal performance in sequential decision-making.
method Proves performance bounds for metalearning priors in stochastic linear bandits and develops a metalearning algorithm.
result Metalearning can improve performance by learning the prior from multiple tasks.
The paper extends and applies a new shrinkage prior in Bayesian factor analysis.
problem Estimating the number of factors in sparse Bayesian factor analysis.
method Introduces and extends a generalized cumulative shrinkage process (CUSP) prior.
result Exchangeable spike-and-slab shrinkage priors imply increasing shrinkage as the column index increases.
Bayesian method corrects for model selection multiplicity in regression.
problem Model selection multiplicity in regression analysis.
method Developed a Bayesian prior distribution based on Holm procedure analogy.
result Adequate multiplicity correction requires sparsity not provided by recommended priors.
Review of priors in Bayesian deep learning models.
problem The importance of prior choices in Bayesian deep learning models.
method Overview of different priors and methods of learning priors from data.
result Motivate practitioners to think carefully about prior specification.
Study characterizes training and test risks for MAP regression with Gaussian priors.
problem Understanding high-dimensional behavior of regularized linear regression with informative priors.
method Maximum a posteriori (MAP) regression with Gaussian priors, using random matrix theory.
result Closed-form risk formulas reveal the bias-variance-prior tradeoff and explain double descent.
This work tackles the challenge of Bayesian deep learning by proposing a new framework for matching Gaussian process priors with neural network parameters.
problem The challenge of specifying priors over neural network parameters, which affects the induced functional prior and is uncontrolled.
method The approach involves defining functional priors using Gaussian processes and matching these priors with the functional prior of neural networks through the minimization of Wasserstein distance.
result The proposed framework offers systematic performance improvements over alternative priors and approximate Bayesian deep learning approaches.
This paper uses reference priors to improve deep learning models with unlabeled and labeled data.
problem Improving deep learning models with limited labeled data and unlabeled data from the same or related tasks.
method Develops and applies generalizations of reference priors for deep networks to exploit unlabeled and labeled data.
result Demonstrates new semi-supervised learning and pretraining methods for transfer learning.
Proposes a new prior for complex models to improve prediction accuracy.
problem Difficulty in specifying priors for complex models like neural networks.
method Predictive complexity priors defined by comparing model predictions to a reference model, transferred to parameters via change of variables.
result Improves model predictions by reducing unintuitive effects of traditional priors.
We construct geometric shrinkage priors for Kählerian signal filters. Based on the characteristics of Kähler manifolds, an efficient and robust algorithm for finding superharmonic priors which outperform the Jeffreys prior is introduced. Several ansätze for the Bayesian predictive priors are also suggested. In particul…
Proposes NUV priors for half-space and box constraints.
problem Adding constraints to linear Gaussian models without computational cost.
method Introduces NUV representations for half-space and box constraints.
result Adds constraints to linear Gaussian models without affecting computational tractability.
Statsformer validates and adapts LLM-derived semantic priors for improved supervised learning.
problem Unreliable semantic priors from LLMs can degrade supervised learning performance.
method Adapts LLM-derived feature scores into a family of learner-specific prior-injection mechanisms, calibrating their influence using out-of-fold validation.
result Improves prediction performance by adaptively downweighting unreliable LLM priors, ensuring a guardrailed statistical learning system.
New priors can update posteriors without re-estimating likelihoods.
problem Degradation of classification approaches when class priors change.
method Recompute posteriors using recovered likelihoods from original posteriors and new priors.
result Dynamic update of original posteriors is possible without re-estimating likelihoods.
SAHMM-VAE separates sources adaptively using hidden Markov priors.
problem Unsupervised blind source separation.
method Source-wise adaptive Hidden Markov prior variational autoencoder.
result Different latent dimensions align with different source-specific temporal organizations.
BNNpriors library improves Bayesian neural network inference with various prior distributions.
problem Challenges in choosing good prior distributions for Bayesian neural networks.
method State-of-the-art Markov Chain Monte Carlo inference with a wide range of predefined priors.
result Facilitates foundational discoveries on the nature of the cold posterior effect.
GOAT improves attention mechanisms by learning better priors.
problem Standard attention mechanisms use a naive uniform prior, limiting flexibility and generalization.
method GOAT introduces a trainable, continuous prior that replaces the uniform assumption, maintaining compatibility with optimized kernels.
result GOAT avoids representational trade-offs and learns an extrapolatable prior that combines positional flexibility with length generalization.
We study the problem of learning shared structure \emph{across} a sequence of dynamic pricing experiments for related products. We consider a practical formulation where the unknown demand parameters for each product come from an unknown distribution (prior) that is shared across products. We then propose a meta dynami…
Weak diffusion priors can still perform well in inverse problems.
problem Using mismatched or low-fidelity diffusion priors in inverse problems.
method Extensive experiments and theoretical analysis combining Bayesian-consistency theory and local-correlation analysis.
result Weak priors succeed when measurements are highly informative, and they fail in other regimes.
New method samples Jeffreys prior for objective Bayesian inference.
problem Sampling from Jeffreys prior is challenging.
method Metropolis-Adjusted Langevin Algorithm
result Samples can be directly used in Bayesian methods.
Optimality of TS with noninformative priors proven for Pareto model.
problem Optimality of Thompson Sampling with noninformative priors for Pareto bandits.
method Proved optimality of TS with certain probability matching priors, showed suboptimality with others, and found effectiveness of truncation procedures.
result TS with certain probability matching priors achieves optimal regret bound for Pareto model.