New confidence intervals improve treatment effect estimation in randomized experiments.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We provide a nonasymptotic analysis of the convergence of the stochastic gradient Hamiltonian Monte Carlo (SGHMC) to a target measure in Wasserstein-2 distance without assuming log-concavity. Our analysis quantifies key theoretical properties of the SGHMC as a sampler under local conditions which significantly improves…
We establish the first nonasymptotic error bounds for Kaplan-Meier-based nearest neighbor and kernel survival probability estimators where feature vectors reside in metric spaces. Our bounds imply rates of strong consistency for these nonparametric estimators and, up to a log factor, match an existing lower bound for c…
ROOT-SGD solves convex optimization problems with optimal nonasymptotic and near-optimal asymptotic performance.
The paper improves nonparametric confidence bands for band-limited functions.
This paper tightens the law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
Optimized AIS scheme reduces bias and MSE for general proposals.
FIEM accelerates EM for large datasets with nonasymptotic convergence bounds.
This paper introduces time-uniform CLT-based confidence intervals for statistical inference.
New method estimates gradients accurately with sharp bounds.
In this work, we study the problem of reconstructing shapes from simple nonasymptotic densities measured only along shape boundaries. The particular density we study is also known as the integral area invariant and corresponds to the area of a disk centered on the boundary that is also inside the shape. It is easy to s…
We propose a new algorithm---Stochastic Proximal Langevin Algorithm (SPLA)---for sampling from a log concave distribution. Our method is a generalization of the Langevin algorithm to potentials expressed as the sum of one stochastic smooth term and multiple stochastic nonsmooth terms. In each iteration, our splitting t…
Gradient descent optimally trains RNNs without overparameterization.
Stochastic Gradient Langevin Dynamics (SGLD) is a popular variant of Stochastic Gradient Descent, where properly scaled isotropic Gaussian noise is added to an unbiased estimate of the gradient at each iteration. This modest change allows SGLD to escape local minima and suffices to guarantee asymptotic convergence to g…
The paper creates nonparametric confidence bands for band-limited functions.
Study on signal detection in sparse additive models with nonasymptotic minimax rates.
Paper improves confidence intervals and variance estimation for deep learning models.
Sampling from various kinds of distributions is an issue of paramount importance in statistics since it is often the key ingredient for constructing estimators, test procedures or confidence intervals. In many situations, the exact sampling from a given distribution is impossible or computationally expensive and, there…
New framework robustifies loss functions with quantiles for outlier resistance.
New theory explains how overparametrized neural networks generalize well without bias-variance trade-off.
Study shows generative priors improve rank-one matrix recovery with optimal sample complexity.
Develops a new algorithm for estimating model parameters using interacting particle systems.
This research provides theoretical guarantees for hyperparameter estimation in complex network dynamical systems.
The paper develops tests for comparing means in high dimensions with unknown covariance.
We study the problem of robustly estimating the posterior distribution for the setting where observed data can be contaminated with potentially adversarial outliers. We propose Rob-ULA, a robust variant of the Unadjusted Langevin Algorithm (ULA), and provide a finite-sample analysis of its sampling distribution. In par…
We consider the problem of high-dimensional Ising (graphical) model selection. We propose a simple algorithm for structure estimation based on the thresholding of the empirical conditional variation distances. We introduce a novel criterion for tractable graph families, where this method is efficient, based on the pres…
Modern geometric measure theory, developed largely to solve the Plateau problem, has generated a great deal of technical machinery which is unfortunately regarded as inaccessible by outsiders. Some of its tools (e.g., flat norm distance and decomposition in generalized surface space) hold interest from a theoretical pe…
The paper provides bounds for regression schemes using nonstationary training samples.
In this paper, we are concerned with a non-asymptotic analysis of sampling algorithms used in nonconvex optimization. In particular, we obtain non-asymptotic estimates in Wasserstein-1 and Wasserstein-2 distances for a popular class of algorithms called Stochastic Gradient Langevin Dynamics (SGLD). In addition, the afo…
We take a Hamiltonian-based perspective to generalize Nesterov's accelerated gradient descent and Polyak's heavy ball method to a broad class of momentum methods in the setting of (possibly) constrained minimization in Euclidean and non-Euclidean normed vector spaces. Our perspective leads to a generic and unifying non…
We give a complete characterization of the sampling complexity of best Markovian arm identification in one-parameter Markovian bandit models. We derive instance specific nonasymptotic and asymptotic lower bounds which generalize those of the IID setting. We analyze the Track-and-Stop strategy, initially proposed for th…
Langevin Monte Carlo (LMC) is an iterative algorithm used to generate samples from a distribution that is known only up to a normalizing constant. The nonasymptotic dependence of its mixing time on the dimension and target accuracy is understood mainly in the setting of smooth (gradient-Lipschitz) log-densities, a seri…
The paper explains how nearest neighbor methods succeed in prediction.
In this paper, we explore a general Aggregated Gradient Langevin Dynamics framework (AGLD) for the Markov Chain Monte Carlo (MCMC) sampling. We investigate the nonasymptotic convergence of AGLD with a unified analysis for different data accessing (e.g. random access, cyclic access and random reshuffle) and snapshot upd…
Efficiently transforms samples from various statistical models.
Deep networks can perfectly classify two low-dimensional manifolds on a sphere with large depth and width.
ConquerNet smooths quantile regression for deep learning with minimax guarantees.
Phase retrieval refers to the problem of recovering real- or complex-valued vectors from magnitude measurements. The best-known algorithms for this problem are iterative in nature and rely on so-called spectral initializers that provide accurate initialization vectors. We propose a novel class of estimators suitable fo…
The Rasch model is widely used for item response analysis in applications ranging from recommender systems to psychology, education, and finance. While a number of estimators have been proposed for the Rasch model over the last decades, the available analytical performance guarantees are mostly asymptotic. This paper p…
New inequalities for matrix supermartingales converge under various conditions.
Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence is known to be fragile. To understand the instability of actor-critic, we focus o…
Range penalization enhances statistical accuracy and resource efficiency in federated learning.
New method approximates sampling from smooth potential distributions using a vanishing penalty.
The paper shows how label noise in training can lead to solutions that solve a Lasso program.
Sparse principal component analysis (PCA) is an important technique for dimensionality reduction of high-dimensional data. However, most existing sparse PCA algorithms are based on non-convex optimization, which provide little guarantee on the global convergence. Sparse PCA algorithms based on a convex formulation, for…
Geometric formalism views optimization algorithms as discrete connections, revealing their algebraic curvature and flatness properties.
We study sparse principal components analysis in high dimensions, where (the number of variables) can be much larger than (the number of observations), and analyze the problem of estimating the subspace spanned by the principal eigenvectors of the population covariance matrix. We introduce two complementary not…
A new method detects changes in data sequences by comparing backward and forward confidence sequences.