Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

152305457609 · Jun 202019922001200920172026
48 results for Sampling Theorems

New findings on boosting sample complexity and implications for hardcore theorem.

problem Understanding the sample complexity of smooth boosting and its implications.
method Analyzing the sample complexity of smooth boosting and relating it to the hardcore theorem.
result The sample complexity of smooth boosting matches existing overhead and provides a separation from distribution-independent boosting.

Let S=Γ\HS=Γ\backslash \mathbb{H} be a hyperbolic surface of finite topological type, such that the Fuchsian group ΓPSL2(R)Γ\le \operatorname{PSL}_2(\mathbb{R}) is non-elementary, and consider any generating set S\mathfrak S of ΓΓ. When sampling by an nn-step random walk in π1(S)Γπ_1(S) \cong Γ with each step given by an element…

2018-07-10abs ↗pdf ↗

The paper studies how more data affects prediction risk in high-dimensional models.

problem The impact of increasing data on prediction risk in high-dimensional models.
method Derives central limit theorem and provides finite-sample distribution and confidence interval for prediction risk.
result Demonstrates 'more data hurt' phenomenon in high-dimensional least squares estimation.

DM improves self-supervised transfer learning by matching target distributions.

problem Improving self-supervised transfer learning performance.
method Distribution Matching (DM) method that drives representation distribution towards a predefined reference distribution.
result DM outperforms existing methods on target classification tasks.

Relationships that exist between the classical, Shannon-type, and geometric-based approaches to sampling are investigated. Some aspects of coding and communication through a Gaussian channel are considered. In particular, a constructive method to determine the quantizing dimension in Zador's theorem is provided. A geom…

2010-02-15abs ↗pdf ↗

A theorem for debiasing machine learning with finite sample guarantees.

problem Calculating confidence intervals for machine learning functionals.
method Debiased machine learning based on bias correction and sample splitting.
result Nonasymptotic debiased machine learning theorem with finite sample guarantees.

We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error pr…

2019-08-27abs ↗pdf ↗

PARIS reduces imbalanced regression datasets by pruning uninformative samples.

problem Imbalanced regression where models focus on high-frequency regions, ignoring rare but impactful events.
method PARIS uses the representer theorem to compute a closed-form representer deletion residual for iterative pruning of the training set.
result PARIS reduces training set by up to 75% while preserving or improving overall performance, outperforming other methods.

The study tightens bounds on binomial probabilities and minimums using KL-divergence.

problem Tightening bounds on binomial probabilities and minimums of i.i.d. Binomials.
method Applied Sanov's theorem to derive upper and lower bounds on binomial tail probabilities and minimums, expressed in terms of KL-divergence.
result High probability upper and lower bounds on the minimum of i.i.d. Binomial random variables, finite sample, asymptotically tight.

Efficient inference method for adaptive experiments with tighter confidence sequences.

problem Efficient inference of Average Treatment Effect in a changing policy sequential experiment.
method Semiparametric efficient inference using Adaptive Augmented Inverse-Probability Weighted estimator and asymptotic confidence sequences.
result Derives tighter confidence sequences for adaptive experiments under data-dependent stopping times.

Develops a new inference method for split-sample estimators using multiple splits.

problem Statistical dependence and variability in split-sample estimators.
method Averaging across multiple splits, proving a central limit theorem, and developing new inference approaches.
result Valid confidence intervals and improved power in comparing model performance.

The paper identifies a 'small' set of functions containing Gaussian process samples.

problem Identifying a small set of functions containing Gaussian process samples.
method Using scaled RKHSs and Karhunen-Loève theorem, the paper defines the sample support set.
result The sample support set consists of functions with bounded squared basis coefficients.

This paper introduces sample-averaged Q-learning for better RL performance.

problem Improving reinforcement learning algorithms by managing uncertainty.
method Integrates statistical inference into Q-learning through sample averaging and functional central limit theorem.
result Establishes a unified theoretical foundation for sample-averaged Q-learning.

The Sampled Gaussian Mechanism's noise level decreases with larger subsampling rates, improving privacy-utility trade-offs.

problem Improving privacy-utility trade-offs in differentially private stochastic optimization.
method Proof of a conjecture about the Sampled Gaussian Mechanism's noise level and subsampling rate relationship.
result A rigorous proof of the conjecture, completing the proof of Theorem 6.2 in the original paper.

UD-SGD analysis shows efficient sampling by a few agents can outperform others.

problem Analyzing convergence speed and sampling strategies in UD-SGD.
method Asymptotic analysis of UD-SGD with various communication patterns and sampling strategies.
result Efficient sampling by a few agents can lead to better overall convergence.

We show that stochastic interpolation flow maps are Lipschitz with a sharp constant.

problem High dimensional sampling and transport problems.
method Investigating stochastic interpolation flow for generating data samples.
result Stochastic interpolation flow maps are Lipschitz with a sharp constant matching optimal transport maps.

Develops an online Gaussian process method that maintains convergence guarantees without sample complexity issues.

problem The computational intractability of Gaussian processes with streaming data.
method Parsimonious Online Gaussian Processes (POG) that maintains asymptotic consistency with bounded memory.
result POG preserves convergence guarantees to the population posterior with finite memory, even for constant error radius.

New theorem improves spectral gap for sampling from mixture distributions.

problem Sampling from multimodal distributions with simulated tempering.
method Introduced a decomposition theorem for the restricted spectral gap of simulated tempering.
result Lower bound on the restricted spectral gap for mixture distributions.

The paper establishes CLTs for Markov chains and improves sampling algorithms for heavy-tailed distributions.

problem Establishing central limit theorems for ergodic averages of Markov chains.
method Drift conditions to provide necessary and sufficient conditions for CLTs, including lower bounds on convergence rates.
result Sharp conditions and convergence rates for various MCMC algorithms on heavy-tailed targets.

This paper strengthens the central limit theorem for order statistics using relative entropy.

problem Establishing a stronger mode of convergence for central limit behavior of order statistics.
method Using relative entropy to ensure a stronger mode of convergence for central limit behavior of order statistics.
result An order O(1/n)O(1/\sqrt{n}) rate of convergence is established under mild conditions.

We propose a general yet simple theorem describing the convergence of SGD under the arbitrary sampling paradigm. Our theorem describes the convergence of an infinite array of variants of SGD, each of which is associated with a specific probability law governing the data selection rule used to form mini-batches. This is…

2019-01-27abs ↗pdf ↗

Study shows a central limit theorem for random coverings of manifolds with nilpotent groups.

problem Understanding the distribution of connected components in random coverings of manifolds with nilpotent fundamental groups.
method Used sampling homomorphisms from the fundamental group into the symmetric group and subgroup growth zeta functions of nilpotent groups.
result Proved a central limit theorem for the number of connected components of these random coverings.

A new approach models exploration in continuous-time RL using random measures.

problem Modeling exploration in continuous-time reinforcement learning.
method Random measure approach to control execution in continuous-time RL.
result Grid-sampling limit SDE can replace existing models for theoretical analysis and learning algorithms.

We consider the problem of estimating E[f(U1,,Ud)]\mathbb{E} [f(U^1, \ldots, U^d)], where (U1,,Ud)(U^1, \ldots, U^d) denotes a random vector with uniformly distributed marginals. In general, Latin hypercube sampling (LHS) is a powerful tool for solving this kind of high-dimensional numerical integration problem. In the case of depende…

2013-11-19abs ↗pdf ↗

Over the past two decades, several consistent procedures have been designed to infer causal conclusions from observational data. We prove that if the true causal network might be an arbitrary, linear Gaussian network or a discrete Bayes network, then every unambiguous causal conclusion produced by a consistent method f…

2012-03-15abs ↗pdf ↗

Conventional approaches of sampling signals follow the celebrated theorem of Nyquist and Shannon. Compressive sampling, introduced by Donoho, Romberg and Tao, is a new paradigm that goes against the conventional methods in data acquisition and provides a way of recovering signals using fewer samples than the traditiona…

2014-05-21abs ↗pdf ↗

The paper strengthens the classical result of MLE convergence to a Gaussian distribution.

problem The classical result of MLE convergence to a Gaussian distribution.
method Sub-Gaussian concentration and entropic normality of the normalized MLE.
result Entropic central limit theorem for a smoothed version of the estimator.

This paper studies binary classification problem associated with a family of loss functions called large-margin unified machines (LUM), which offers a natural bridge between distribution-based likelihood approaches and margin-based approaches. It also can overcome the so-called data piling issue of support vector machi…

2019-08-13abs ↗pdf ↗

Paper presents a new policy gradient theorem using weak derivatives for reinforcement learning.

problem Continuous state-action reinforcement learning problems.
method Introduced an alternative policy gradient theorem using weak derivatives.
result The new approach yields algorithms that converge almost surely to stationary points of the value function.

New analysis improves sample complexity for vanilla policy gradient methods.

problem Improving sample complexity guarantees for vanilla policy gradient methods.
method Adapting tools from SGD analysis to policy gradient methods, with smoothness and gradient approximation assumptions.
result Established improved sample complexity bounds for convergence and global optimum.

Paper introduces a neural network training algorithm for noisy data that achieves optimal parameters and replicates real-world behaviors.

problem Theoretical gap between universal approximation theorems and practical machine learning with noisy data.
method Randomized training algorithm for neural networks trained on noisy data samples.
result Trained neural networks achieve optimal parameters and exhibit real-world behaviors like sub-linear complexity and interpolation.

The study tightens the sample complexity for learning nonparametric mixture components.

problem Learning nonparametric distributions in a finite mixture model.
method Assumes each component is a convolution of a Gaussian and a compactly supported density, and uses a quantitative Tauberian theorem.
result Tight bounds on sample complexity required for estimating each component, showing it lies between polynomial and exponential.

New insights into variational inference using Monte Carlo estimates.

problem Improving variational bounds in latent variable models.
method Analyzing properties of Monte Carlo estimates and their impact on variational gaps.
result Negative correlation reduces variational gaps, contrary to intuition.