Motivated by a range of applications in engineering and genomics, we consider in this paper detection of very short signal segments in three settings: signals with known shape, arbitrary signals, and smooth signals. Optimal rates of detection are established for the three cases and rate-optimal detectors are constructe…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study a nonparametric contextual bandit problem where the expected reward functions belong to a Hölder class with smoothness parameter . We show how this interpolates between two extremes that were previously studied in isolation: non-differentiable bandits (), where rate-optimal regret is achieved by run…
CARROT optimizes LLM routing by choosing the cheapest and most accurate model.
New algorithm optimizes best arm identification with minimal regret.
ERM and RERM minimize error even with malicious label corruptions.
The contextual bandit literature has traditionally focused on algorithms that address the exploration-exploitation tradeoff. In particular, greedy algorithms that exploit current estimates without any exploration may be sub-optimal in general. However, exploration-free greedy algorithms are desirable in practical setti…
Optimally tackles covariate shift in RKHS-based nonparametric regression.
Meta learning of optimal classifier error rates allows an experimenter to empirically estimate the intrinsic ability of any estimator to discriminate between two populations, circumventing the difficult problem of estimating the optimal Bayes classifier. To this end we propose a weighted nearest neighbor (WNN) graph es…
Paper eliminates warm-up phase for PO in linear MDPs, achieving optimal regret.
Paper analyzes minimax risks of personalized federated learning algorithms.
Biclustering structures in data matrices were first formalized in a seminal paper by John Hartigan (1972) where one seeks to cluster cases and variables simultaneously. Such structures are also prevalent in block modeling of networks. In this paper, we develop a unified theory for the estimation and completion of matri…
Optimizes prediction error method for time-varying models.
Study minimax off-policy evaluation in multi-armed bandits with known and unknown behavior policies.
Develops a high-dimensional differentially-private EM algorithm with near-optimal statistical guarantees.
Unintended effects from scaling neural network outputs with adaptive learning rates.
VAV method optimizes learning rate for faster, stable SGD convergence.
New method finds linear relationships across multiple data blocks using proximal gradient descent with constraint.
Estimates treatment effects in panel data with general intervention patterns.
Study on estimating volatility of volatility using Fourier methods and provides insights into volatility dynamics.
A new test statistic speeds up MMD while maintaining power.
Neural networks estimate statistical divergences with performance guarantees.
New methods improve estimation accuracy in noisy settings.
This work provides a simplified proof of the statistical minimax optimality of (iterate averaged) stochastic gradient descent (SGD), for the special case of least squares. This result is obtained by analyzing SGD as a stochastic process and by sharply characterizing the stationary covariance matrix of this process. The…
The basic model for high-frequency data in finance is considered, where an efficient price process is observed under microstructure noise. It is shown that this nonparametric model is in Le Cam's sense asymptotically equivalent to a Gaussian shift experiment in terms of the square root of the volatility function . A…
Study sharp convergence rates of empirical UOT for spatio-temporal point processes.
Needlets have been recognized as state-of-the-art tools to tackle spherical data, due to their excellent localization properties in both spacial and frequency domains. This paper considers developing kernel methods associated with the needlet kernel for nonparametric regression problems whose predictor variables are de…
AEW estimator achieves optimal risk in expectation for large enough temperatures.
Paper analyzes kNN estimator for KL divergence, proving its optimality.
New model for pairwise comparisons without stochastic transitivity.
Greedy algorithms outperform UCB in many-armed bandit problems.
Study optimizes shared singular subspace estimation from noisy matrices.
Study approximates unknown function levels with queries.
We consider streaming principal component analysis when the stochastic data-generating model is subject to perturbations. While existing models assume a fixed covariance, we adopt a robust perspective where the covariance matrix belongs to a temporal uncertainty set. Under this setting, we provide fundamental limits on…
New optimizers improve stock market forecasting accuracy.
We analyze the Kozachenko--Leonenko (KL) nearest neighbor estimator for the differential entropy. We obtain the first uniform upper bound on its performance over Hölder balls on a torus without assuming any conditions on how close the density could be from zero. Accompanying a new minimax lower bound over the Hölder ba…
Kernel ridge regression (KRR) is a well-known and popular nonparametric regression approach with many desirable properties, including minimax rate-optimality in estimating functions that belong to common reproducing kernel Hilbert spaces (RKHS). The approach, however, is computationally intensive for large data sets, d…
We present a convergence rate analysis for biased stochastic gradient descent (SGD), where individual gradient updates are corrupted by computation errors. We develop stochastic quadratic constraints to formulate a small linear matrix inequality (LMI) whose feasible points lead to convergence bounds of biased SGD. Base…
We study estimation of (semi-)inner products between two nonparametric probability distributions, given IID samples from each distribution. These products include relatively well-studied classical and Sobolev inner products, as well as those induced by translation-invariant reproducing kernels, for whic…
New method clusters matrix-valued data by latent variables.
Study optimizes dividend payout strategies under fluctuating interest rates.
In this paper, we study the multi-armed bandit problem in the batched setting where the employed policy must split data into a small number of batches. While the minimax regret for the two-armed stochastic bandits has been completely characterized in \cite{perchet2016batched}, the effect of the number of arms on the re…
We provide upper bounds of the expected Wasserstein distance between a probability measure and its empirical version, generalizing recent results for finite dimensional Euclidean spaces and bounded functional spaces. Such a generalization can cover Euclidean spaces with large dimensionality, with the optimal dependence…
IDS algorithm optimizes sequential decisions in various monitoring settings.
In this paper, we propose a general framework for sparse and low-rank tensor estimation from cubic sketchings. A two-stage non-convex implementation is developed based on sparse tensor decomposition and thresholded gradient descent, which ensures exact recovery in the noiseless case and stable recovery in the noisy cas…
MARTHE optimizes learning rates online using hypergradient approximations.
D-Adaptation automatically sets optimal learning rates without manual tuning.
Optimal control in changing systems without strong convexity assumptions.
We establish the consistency of an algorithm of Mondrian Forests, a randomized classification algorithm that can be implemented online. First, we amend the original Mondrian Forest algorithm, that considers a fixed lifetime parameter. Indeed, the fact that this parameter is fixed hinders the statistical consistency of …