Paper tackles unknown variances in best-arm identification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithms improve best-arm identification with varying rewards.
Sharp inequalities for matrix means with unknown variance.
Algorithm estimates common mean from Gaussian variables with unknown variances.
New strategy optimally identifies best arm in unknown variance Gaussian bandits.
The paper analyzes sparse high-dimensional linear regression with random design and unknown error variance, providing adaptiveness and concentration rates.
Paper presents a method to reduce prediction variance of DNNs for unknown systems.
A new estimator for evaluating policies in unknown environments.
New method optimizes portfolio weights as functions, outperforming traditional approaches.
This paper introduces the first asymptotically optimal strategy for a multi armed bandit (MAB) model under side constraints. The side constraints model situations in which bandit activations are limited by the availability of certain resources that are replenished at a constant rate. The main result involves the deriva…
New algorithm reduces regret for linear bandits with unknown noise variance.
Develops new e-processes and confidence sequences for Gaussian means with unknown variance.
The paper proposes a method for distribution-free prediction sets that adapt to unknown temporal changes.
Improved SGD with AdaGrad stepsizes adapts to unknown parameters and unbounded gradients.
We study confidence intervals based on hard-thresholding, soft-thresholding, and adaptive soft-thresholding in a linear regression model where the number of regressors may depend on and diverge with sample size . In addition to the case of known error variance, we define and study versions of the estimators when…
The paper develops adaptive confidence intervals for Efron's Gaussian two-groups model with unknown contamination.
This paper investigates methods for estimating the optimal stochastic control policy for a Markov Decision Process with unknown transition dynamics and an unknown reward function. This form of model-free reinforcement learning comprises many real world systems such as playing video games, simulated control tasks, and r…
Thompson sampling used for linear bandits with normal-gamma priors.
Existing strategies for finite-armed stochastic bandits mostly depend on a parameter of scale that must be known in advance. Sometimes this is in the form of a bound on the payoffs, or the knowledge of a variance or subgaussian parameter. The notable exceptions are the analysis of Gaussian bandits with unknown mean and…
Optimizes budgeted evaluations of LLMs by allocating queries to judges efficiently.
Optimal B-robust estimate is constructed for multidimensional parameter in drift coefficient of diffusion type process with small noise. Optimal mean-variance robust (optimal V -robust) trading strategy is find to hedge in mean-variance sense the contingent claim in incomplete financial market with arbitrary informatio…
New algorithms reduce contextual bandits' regret without knowing reward noise variances.
Markowitz' celebrated optimal portfolio theory generally fails to deliver out-of-sample diversification. In this note, we propose a new portfolio construction strategy based on symmetry arguments only, leading to "Eigenrisk Parity" portfolios that achieve equal realized risk on all the principal components of the covar…
Gradient-based methods for optimisation of objectives in stochastic settings with unknown or intractable dynamics require estimators of derivatives. We derive an objective that, under automatic differentiation, produces low-variance unbiased estimators of derivatives at any order. Our objective is compatible with arbit…
NP-PROV separates mean and variance spaces to improve function uncertainty.
The paper estimates common mean of entangled Gaussians with bounded variances.
Bayesian investor learns unknown asset drift, trades mean-variance optimal portfolio, but policy is robust to observation model distortion.
Markowitz's celebrated mean--variance portfolio optimization theory assumes that the means and covariances of the underlying asset returns are known. In practice, they are unknown and have to be estimated from historical data. Plugging the estimates into the efficient frontier that assumes known parameters has led to p…
When randomized ensembles such as bagging or random forests are used for binary classification, the prediction error of the ensemble tends to decrease and stabilize as the number of classifiers increases. However, the precise relationship between prediction error and ensemble size is unknown in practice. In the standar…
Variance reduction is a simple and effective technique that accelerates convex (or non-convex) stochastic optimization. Among existing variance reduction methods, SVRG and SAGA adopt unbiased gradient estimators and are the most popular variance reduction methods in recent years. Although various accelerated variants o…
We address the issue of estimating the regression vector in the generic -sparse linear model , with , , $z\sim\mathcal N(0,\sg^2 I)$ and when the variance $\sg^{2}$ is unknown. We study two LASSO-type methods that jointly estimate and the variance. These estimators ar…
Consider the problem of sampling sequentially from a finite number of populations, specified by random variables , and ; where denotes the outcome from population the time it is sampled. It is assumed that for each fixed , $\{ X^i_k \}_{k …
In the paper, we consider three quadratic optimization problems which are frequently applied in portfolio theory, i.e, the Markowitz mean-variance problem as well as the problems based on the mean-variance utility function and the quadratic utility.Conditions are derived under which the solutions of these three optimiz…
Efficient RL for linear MDPs with unknown transitions.
We propose a Bayesian expectation-maximization (EM) algorithm for reconstructing Markov-tree sparse signals via belief propagation. The measurements follow an underdetermined linear model where the regression-coefficient vector is the sum of an unknown approximately sparse signal and a zero-mean white Gaussian noise wi…
Improved GP bandit algorithms for noiseless, varying noise, and RKHS norms.
For many important problems the quantity of interest is an unknown function of the parameters, which is a random vector with known statistics. Since the dependence of the output on this random vector is unknown, the challenge is to identify its statistics, using the minimum number of function evaluations. This problem …
Improved mean estimation for symmetric distributions with finite-sample guarantees.
Faster convergence of kernel mean embeddings using variance information.
We study the problem of estimating low-rank matrices from linear measurements (a.k.a., matrix sensing) through nonconvex optimization. We propose an efficient stochastic variance reduced gradient descent algorithm to solve a nonconvex optimization problem of matrix sensing. Our algorithm is applicable to both noisy and…
Variational Bayes (VB) is a recent approximate method for Bayesian inference. It has the merit of being a fast and scalable alternative to Markov Chain Monte Carlo (MCMC) but its approximation error is often unknown. In this paper, we derive the approximation error of VB in terms of mean, mode, variance, predictive den…
The paper develops a robust algorithm for contextual bandits with heavy-tailed rewards.
UCB-V algorithm improves on UCB for MAB problems with variance estimates.
New algorithm tackles adversarial bandits with arbitrary strategies.
This paper addresses error bounds and posterior variance for Gaussian process regression.
Method estimates group structure in panel data using variance information.
In markets for online advertising, some advertisers pay only when users respond to ads. So publishers estimate ad response rates and multiply by advertiser bids to estimate expected revenue for showing ads. Since these estimates may be inaccurate, the publisher risks not selecting the ad for each ad call that would max…
We study a distributed estimation problem in which two remotely located parties, Alice and Bob, observe an unlimited number of i.i.d. samples corresponding to two different parts of a random vector. Alice can send bits on average to Bob, who in turn wants to estimate the cross-correlation matrix between the two par…