The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New approach to neural networks by incorporating observation noise and arbitrary prior means.
When fitting Bayesian machine learning models on scarce data, the main challenge is to obtain suitable prior knowledge and encode it into the model. Recent advances in meta-learning offer powerful methods for extracting such prior knowledge from data acquired in related tasks. When it comes to meta-learning in Gaussian…
A new data-adaptive prior stabilizes kernel learning in operators.
It is well-known that the distribution over functions induced through a zero-mean iid prior distribution over the parameters of a multi-layer perceptron (MLP) converges to a Gaussian process (GP), under mild conditions. We extend this result firstly to independent priors with general zero or non-zero means, and secondl…
This paper improves Gaussian process predictions by integrating prior knowledge.
Paper proposes a method to break symmetries in Bayesian matrix factorization.
Study shows prior Lipschitz continuity can improve adversarial robustness of Bayesian Neural Networks.
Additive Bayesian networks are types of graphical models that extend the usual Bayesian generalized linear model to multiple dependent variables through the factorisation of the joint probability distribution of the underlying variables. When fitting an ABN model, the choice of the prior of the parameters is of crucial…
The paper deals with learning probability distributions of observed data by artificial neural networks. We suggest a so-called gradient conjugate prior (GCP) update appropriate for neural networks, which is a modification of the classical Bayesian update for conjugate priors. We establish a connection between the gradi…
The paper improves bandit algorithms by incorporating random-effect models.
We consider classical Merton problem of terminal wealth maximization in finite horizon. We assume that the drift of the stock is following Ornstein-Uhlenbeck process and the volatility of it is following GARCH(1) process. In particular, both mean and volatility are unbounded. We assume that there is Knightian uncertain…
The paper examines conditions for linearity in a conditional mean estimator under vector Poisson noise.
CW-Gen models improve probabilistic time series forecasting by incorporating prior information.
New meta-reinforcement learning method improves performance in finite-horizon MDPs.
New method uses quotient predictor space for better PAC-Bayes bounds, reducing KL divergence and improving model performance.
Bayes-assisted confidence sequences improve efficiency for bounded means.
Flexible Bayesian approach for generalized linear models, especially for sparse logistic regression.
New method uses KL-divergence to create non-informative priors for multivariate Gaussian.
Wide BNNs with odd activations fail to approximate data under mean-field inference.
Beta process is the standard nonparametric Bayesian prior for latent factor model. In this paper, we derive a structured mean-field variational inference algorithm for a beta process non-negative matrix factorization (NMF) model with Poisson likelihood. Unlike the linear Gaussian model, which is well-studied in the non…
New insights into empirical Bayes and compound decision problems with improved regret bounds.
The explore{exploit dilemma is one of the central challenges in Reinforcement Learning (RL). Bayesian RL solves the dilemma by providing the agent with information in the form of a prior distribution over environments; however, full Bayesian planning is intractable. Planning with the mean MDP is a common myopic approxi…
The choice of sentence encoder architecture reflects assumptions about how a sentence's meaning is composed from its constituent words. We examine the contribution of these architectures by holding them randomly initialised and fixed, effectively treating them as as hand-crafted language priors, and evaluating the resu…
The concept of sample mean in dynamic time warping (DTW) spaces has been successfully applied to improve pattern recognition systems and generalize centroid-based clustering algorithms. Its existence has neither been proved nor challenged. This article presents sufficient conditions for existence of a sample mean in DT…
A method for converting NIW parameters for better estimation.
New estimator accurately estimates mean of real-valued distributions without variance knowledge.
Exemplar-based clustering methods have been shown to produce state-of-the-art results on a number of synthetic and real-world clustering problems. They are appealing because they offer computational benefits over latent-mean models and can handle arbitrary pairwise similarity measures between data points. However, when…
We consider the stochastic multi-armed bandit problem with a prior distribution on the reward distributions. We are interested in studying prior-free and prior-dependent regret bounds, very much in the same spirit as the usual distribution-free and distribution-dependent bounds for the non-Bayesian stochastic bandit. B…
New GP kernels avoid mean reversion without losing smoothness.
Bayesian priors offer a compact yet general means of incorporating domain knowledge into many learning tasks. The correctness of the Bayesian analysis and inference, however, largely depends on accuracy and correctness of these priors. PAC-Bayesian methods overcome this problem by providing bounds that hold regardless …
Compressive sensing is an impressive approach for fast MRI. It aims at reconstructing MR image using only a few under-sampled data in k-space, enhancing the efficiency of the data acquisition. In this study, we propose to learn priors based on undecimated wavelet transform and an iterative image reconstruction algorith…
We study convergence rates of variational posterior distributions for nonparametric and high-dimensional inference. We formulate general conditions on prior, likelihood, and variational class that characterize the convergence rates. Under similar "prior mass and testing" conditions considered in the literature, the rat…
Adversarial meta-learning computes Gamma-minimax estimators for vague prior knowledge.
In this paper, we introduce a new sparsity-promoting prior, namely, the "normal product" prior, and develop an efficient algorithm for sparse signal recovery under the Bayesian framework. The normal product distribution is the distribution of a product of two normally distributed variables with zero means and possibly …
Study proposes learning optimal priors from data for better Bayesian inference.
New algorithm clusters data with almost-linear time, robust to corruption.
The Probably Approximately Correct (PAC) Bayes framework (McAllester, 1999) can incorporate knowledge about the learning algorithm and (data) distribution through the use of distribution-dependent priors, yielding tighter generalization bounds on data-dependent posteriors. Using this flexibility, however, is difficult,…
Thompson sampling has impressive empirical performance for many multi-armed bandit problems. But current algorithms for Thompson sampling only work for the case of conjugate priors since these algorithms require to infer the posterior, which is often computationally intractable when the prior is not conjugate. In this …
Variational dropout (VD) is a generalization of Gaussian dropout, which aims at inferring the posterior of network weights based on a log-uniform prior on them to learn these weights as well as dropout rate simultaneously. The log-uniform prior not only interprets the regularization capacity of Gaussian dropout in netw…
Bayesian framework for sphere regression using Gaussian fields.
Study Nash equilibrium in market with relative wealth concerns under partial information and heterogeneous priors.
New data-dependent priors improve PAC-Bayes bounds.
Proposes I-prior extension for additive interaction models.
Deep Gaussian Processes with polynomial kernels can collapse rapidly without proper hyperparameter tuning.
New method for high-dimensional linear regression using empirical Bayes.
Proposes a method to quantify uncertainty in PFNs.
Develops methods for structured variational inference with star-structured models.