Reduces quantifier variance with accuracy optimization of base classifier.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Thompson sampling used for linear bandits with normal-gamma priors.
High-dimensional shrinkage risk depends on the default prior for the common scale.
We study the problem of finding probability densities that match given European call option prices. To allow prior information about such a density to be taken into account, we generalise the algorithm presented in Neri and Schneider (2011) to find the maximum entropy density of an asset price to the relative entropy c…
FlexAE addresses bias-variance trade-off in RAEs by learning latent priors.
A novel Bayesian method for dynamic sparsity in Gaussian dynamic linear regression.
Bayesian framework improves variance component estimation in MET data.
Dealing with high variance is a significant challenge in model-free reinforcement learning (RL). Existing methods are unreliable, exhibiting high variance in performance from run to run using different initializations/seeds. Focusing on problems arising in continuous control, we propose a functional regularization appr…
Study shows prior Lipschitz continuity can improve adversarial robustness of Bayesian Neural Networks.
The paper analyzes sparse high-dimensional linear regression with random design and unknown error variance, providing adaptiveness and concentration rates.
Bayesian neural networks (BNNs) hold great promise as a flexible and principled solution to deal with uncertainty when learning from finite data. Among approaches to realize probabilistic inference in deep neural networks, variational Bayes (VB) is theoretically grounded, generally applicable, and computationally effic…
New estimator accurately estimates mean of real-valued distributions without variance knowledge.
We analyze a new robust method for the reconstruction of probability distributions of observed data in the presence of output outliers. It is based on a so-called gradient conjugate prior (GCP) network which outputs the parameters of a prior. By rigorously studying the dynamics of the GCP learning process, we derive an…
Additive Bayesian networks are types of graphical models that extend the usual Bayesian generalized linear model to multiple dependent variables through the factorisation of the joint probability distribution of the underlying variables. When fitting an ABN model, the choice of the prior of the parameters is of crucial…
The paper deals with learning probability distributions of observed data by artificial neural networks. We suggest a so-called gradient conjugate prior (GCP) update appropriate for neural networks, which is a modification of the classical Bayesian update for conjugate priors. We establish a connection between the gradi…
A trade-off exists between reconstruction quality and the prior regularisation in the Evidence Lower Bound (ELBO) loss that Variational Autoencoder (VAE) models use for learning. There are few satisfactory approaches to deal with a balance between the prior and reconstruction objective, with most methods dealing with t…
Variational dropout (VD) is a generalization of Gaussian dropout, which aims at inferring the posterior of network weights based on a log-uniform prior on them to learn these weights as well as dropout rate simultaneously. The log-uniform prior not only interprets the regularization capacity of Gaussian dropout in netw…
Proposes NUV priors for half-space and box constraints.
We consider learning on graphs, guided by kernels that encode similarity between vertices. Our focus is on random walk kernels, the analogues of squared exponential kernels in Euclidean spaces. We show that on large, locally treelike, graphs these have some counter-intuitive properties, specifically in the limit of lar…
This paper presents a novel approach for approximate integration over the uncertainty of noise and signal variances in Gaussian process (GP) regression. Our efficient and straightforward approach can also be applied to integration over input dependent noise variance (heteroscedasticity) and input dependent signal varia…
Bayesian method recovers causal structure in SEMs with equal error variances.
Generalizes bias-variance decomposition for Bregman divergences.
This work uses ANOVA to understand how different factors contribute to test error in machine learning models.
New algorithm reduces regret in infinite MDPs with optimal variance-dependent bounds.
A simple method treats heteroscedastic variance variatively, improving model calibration and sample quality.
Two new estimators improve VAE training for hierarchical and prior parameters.
Adam converges with high probability under unconstrained non-convex smooth stochastic optimizations.
Statistical physics approaches can be used to derive accurate predictions for the performance of inference methods learning from potentially noisy data, as quantified by the learning curve defined as the average error versus number of training examples. We analyse a challenging problem in the area of non-parametric inf…
The study finds that memorization is necessary or harmful depending on the prior distribution and noise level.
New algorithms improve best-arm identification with varying rewards.
Thermodynamic integration (TI) for computing marginal likelihoods is based on an inverse annealing path from the prior to the posterior distribution. In many cases, the resulting estimator suffers from high variability, which particularly stems from the prior regime. When comparing complex models with differences in a …
MFVI can overestimate predictive variance compared to the exact posterior
Study characterizes training and test risks for MAP regression with Gaussian priors.
Adaptive importance sampling for stochastic optimization is a promising approach that offers improved convergence through variance reduction. In this work, we propose a new framework for variance reduction that enables the use of mixtures over predefined sampling distributions, which can naturally encode prior knowledg…
Enhances neural network regression performance by modeling weight and variance uncertainty.
While Bayesian neural networks have many appealing characteristics, current priors do not easily allow users to specify basic properties such as expected lengthscale or amplitude variance. In this work, we introduce Poisson Process Radial Basis Function Networks, a novel prior that is able to encode amplitude stationar…
Bayesian neural networks (BNNs) have developed into useful tools for probabilistic modelling due to recent advances in variational inference enabling large scale BNNs. However, BNNs remain brittle and hard to train, especially: (1) when using deep architectures consisting of many hidden layers and (2) in situations wit…
We use the P&L on a particular class of swaps, representing variance and higher moments for log returns, as estimators in our empirical study on the S&P500 that investigates the factors determining variance and higher-moment risk premia. This class is the discretisation invariant sub-class of swaps with Neuberger's agg…
Bayesian PROCOVA uses AI to adjust for covariates in RCTs.
IENs reduce neural network variance without increasing complexity.
In this paper, we introduce a new sparsity-promoting prior, namely, the "normal product" prior, and develop an efficient algorithm for sparse signal recovery under the Bayesian framework. The normal product distribution is the distribution of a product of two normally distributed variables with zero means and possibly …
FedGLOMO accelerates FL convergence for non-convex functions.
Improves diffusion models by controlling total variance and signal-to-noise-ratio.
Study Nash equilibrium in market with relative wealth concerns under partial information and heterogeneous priors.
Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the state and action and suggest that this significantly reduces variance and improv…
Student- processes have recently been proposed as an appealing alternative non-parameteric function prior. They feature enhanced flexibility and predictive variance. In this work the use of Student- processes are explored for multi-objective Bayesian optimization. In particular, an analytical expression for the h…
Optimal feature transfer identified through bias-variance analysis.
Encoding domain knowledge into the prior over the high-dimensional weight space of a neural network is challenging but essential in applications with limited data and weak signals. Two types of domain knowledge are commonly available in scientific applications: 1. feature sparsity (fraction of features deemed relevant)…