Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

65129194258 · May 202619922001200920182026
48 results for large-sample theory

Infinitesimal boosting converges to a deterministic process in large sample limit.

problem Characterizing the asymptotic behavior of infinitesimal gradient boosting in large sample sizes.
method Proving convergence to a deterministic process using large sample theory and differential equations.
result The test error decreases over time in the population limit.

Develops large-sample theory for non-stationary source separation.

problem Lack of large-sample results for non-stationary source separation methods.
method Large-sample theory for NSS-JD method under specific assumptions.
result Consistency of unmixing estimator and its convergence to Gaussian distribution.

We analyze SGAs for statistical inference via asymptotics, improving tuning methods.

problem Improper tuning of SGAs for optimization and sampling.
method Characterize large-sample asymptotics of SGAs via step-size and sample-size scaling limits.
result Iterate averaging with large step size is robust and asymptotically has covariance proportional to MLE's.

New method compresses large sample data for faster discriminant analysis.

problem Large sample sizes in discriminant analysis increase computational burden.
method Proposes a new compression approach for reducing training samples.
result Significant computational gains and superior predictive ability compared to random sub-sampling.

New framework analyzes SGD dynamics in large samples and dimensions.

problem Analyzing stochastic gradient descent in large-scale settings.
method Inspired by random matrix theory, new framework for fixed stepsize and finite sum settings.
result SGD dynamics become deterministic in the large sample and dimensional limit, governed by a Volterra integral equation.

Develops stability conditions for estimating affine jump-diffusions.

problem Ergodicity and consistency of parameter estimation for affine jump-diffusions.
method Establishes stochastic stability conditions and ergodicity under specific conditions.
result Proves strong laws of large numbers and functional central limit theorems for additive functionals.

For a finite function class we describe the large sample limit of the sequential Rademacher complexity in terms of the viscosity solution of a GG-heat equation. In the language of Peng's sublinear expectation theory, the same quantity equals to the expected value of the largest order statistics of a multidimensional $…

2016-05-11abs ↗pdf ↗

New tuning rules for Metropolis algorithms derived from Bayesian large-sample asymptotics.

problem Optimal scaling in random-walk Metropolis algorithms under realistic assumptions.
method Large-sample asymptotics to derive weak convergence results and tuning guidelines.
result Tuning guidelines consistent with previous ones when target density is product form, accounting for correlation structure.

New auditors assess ff-DP privacy with adaptive sampling, avoiding large sample sizes.

problem Empirical auditing of ff-DP privacy with adaptive sampling.
method Shift focus to ff-DP, develop adaptive auditors for whitebox and blackbox settings.
result Adaptive auditors detect ff-DP violations across the privacy spectrum with statistical guarantees.

We study the distribution of the adaptive LASSO estimator (Zou (2006)) in finite samples as well as in the large-sample limit. The large-sample distributions are derived both for the case where the adaptive LASSO estimator is tuned to perform conservative model selection as well as for the case where the tuning results…

2008-01-30abs ↗pdf ↗

Two new methods for analyzing repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.

problem Analyzing complex data structures with multiple features over time.
method Two generalizations of canonical correlation analysis for repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.
result Consistency rates for transformation and correlation estimators, relaxing common assumptions.

New methods for estimating causal effects with limited overlap, using Stable Probability Weighting.

problem Estimating causal effects with limited overlap in multivalued treatments.
method Stable Probability Weighting (SPW) and Finite-Sample Stable Probability Weighting (FPW) methods.
result SPW and FPW provide practical solutions for estimating and inferring causal effects with limited overlap.

The paper analyzes LIME for tabular data and proves its behavior in large samples.

problem Understanding the behavior of LIME in tabular data settings.
method Theoretical analysis of LIME's behavior in tabular data, proving its properties in the large sample limit.
result LIME provides explanations proportional to the coefficients of the function in linear cases, but can produce misleading explanations for partition-based models.

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample from the original full sample and uses it as a surrogate for subsequent computati…

2015-09-17abs ↗pdf ↗

Although consistency is a minimum requirement of any estimator, little is known about consistency of the mean partition approach in consensus clustering. This contribution studies the asymptotic behavior of mean partitions. We show that under normal assumptions, the mean partition approach is consistent and asymptotic …

2015-12-18abs ↗pdf ↗

New methods estimate interventional effects with multiple mediators using machine learning.

problem Estimating interventional effects with multiple mediators.
method Flexible machine learning techniques for estimation, with weak convergence results for confidence intervals.
result Closed-form confidence intervals and hypothesis tests for interventional mediation effects.

A practical algorithm improves approximate OT distances using quantization.

problem Substantial computational burden in computing OT distances for large samples.
method Introduces a quantization step to estimate OT distances between measures.
result The quantization step improves the performance of approximate solvers for entropy-regularized transport.

This paper introduces online algorithms to estimate robust geometric median in large data streams.

problem Detecting outliers in large data sets using robust statistical measures.
method Online stochastic Newton methods for estimating the geometric median.
result Rates of convergence for online estimation of the geometric median.

We apply random matrix theory to derive spectral density of large sample covariance matrices generated by multivariate VMA(q), VAR(q) and VARMA(q1,q2) processes. In particular, we consider a limit where the number of random variables N and the number of consecutive time measurements T are large but the ratio N/T is fix…

2010-02-04abs ↗pdf ↗

In kernel methods, the median heuristic has been widely used as a way of setting the bandwidth of RBF kernels. While its empirical performances make it a safe choice under many circumstances, there is little theoretical understanding of why this is the case. Our aim in this paper is to advance our understanding of the …

2017-07-23abs ↗pdf ↗

The paper proves asymptotic normality for multinomial logistic regression on null covariates.

problem Classical asymptotic normality results fail in high-dimensional multinomial logistic models.
method Developed asymptotic normality and chi-square results for multinomial logistic MLE on null covariates.
result Validated new methodology to test feature significance in high-dimensional classification problems.

We theoretically discuss why deep neural networks (DNNs) performs better than other models in some cases by investigating statistical properties of DNNs for non-smooth functions. While DNNs have empirically shown higher performance than other standard methods, understanding its mechanism is still a challenging problem.…

2018-02-13abs ↗pdf ↗

For spherically symmetric distributions, efficient quantisation can be achieved with moderate sample sizes.

problem Optimal quantisation in high dimensions requires large sample sizes, making it impractical.
method Uniformly distributed random quantisers on a sphere of suitable radius achieve exceptional performance.
result For moderate sample sizes, quantisation error can be efficiently computed and approximated.

Maximum Variance Unfolding is one of the main methods for (nonlinear) dimensionality reduction. We study its large sample limit, providing specific rates of convergence under standard assumptions. We find that it is consistent when the underlying submanifold is isometric to a convex subset, and we provide some simple e…

2012-08-31abs ↗pdf ↗

New method improves uncertainty quantification for large batch sizes and misspecified models.

problem Challenges in tuning algorithms for accurate uncertainty quantification in large batch sizes and misspecified models.
method Proposes new discrete-time approximations to SGD and SGLD, proving error bounds for practical tuning.
result Quantitative, non-asymptotic error bounds for accurate predictions of covariance and autocorrelation time.

Online (also called "recursive" or "adaptive") estimation of fixed model parameters in hidden Markov models is a topic of much interest in times series modelling. In this work, we propose an online parameter estimation algorithm that combines two key ideas. The first one, which is deeply rooted in the Expectation-Maxim…

2009-08-17abs ↗pdf ↗

The paper develops a statistical theory explaining overfitting in imbalanced classification.

problem Overfitting in high-dimensional imbalanced classification.
method Developed a statistical theory for support vector machines and logistic regression.
result Overfitting is more severe for the minority class due to truncation or skewing effects in high-dimensional data.

Motivated by safety-critical applications, test-time attacks on classifiers via adversarial examples has recently received a great deal of attention. However, there is a general lack of understanding on why adversarial examples arise; whether they originate due to inherent properties of data or due to lack of training …

2017-06-13abs ↗pdf ↗

We present a new package in R implementing Bayesian additive regression trees (BART). The package introduces many new features for data analysis using BART such as variable selection, interaction detection, model diagnostic plots, incorporation of missing data and the ability to save trees for future prediction. It is …

2013-12-08abs ↗pdf ↗

The question of how to determine the number of independent latent factors (topics) in mixture models such as Latent Dirichlet Allocation (LDA) is of great practical importance. In most applications, the exact number of topics is unknown, and depends on the application and the size of the data set. Bayesian nonparametri…

2013-12-10abs ↗pdf ↗

Investments with best performance are not associated with best Sharpe ratios.

problem The relationship between performance and risk-adjusted return (Sharpe ratio) is counterintuitive for heavy-tailed distributions.
method Synthetic and real data analysis of returns distributions.
result The best-performing investments are not the best in terms of Sharpe ratio, and vice versa.

Study estimates heterogeneous principal causal effects with binary treatments and intermediate variables.

problem Estimating subgroup effects within strata defined by potential values of an intermediate variable.
method Proposes a framework for estimating and forming confidence intervals for heterogeneous principal causal effects under principal ignorability assumption. Develops several estimators with varying robustness properties.
result Established large-sample theory and analyzed bias contributions of each approach.

We propose generalized random forests, a method for non-parametric statistical estimation based on random forests (Breiman, 2001) that can be used to fit any quantity of interest identified as the solution to a set of local moment equations. Following the literature on local maximum likelihood estimation, our method co…

2016-10-05abs ↗pdf ↗

This paper develops DRO estimators for EVT statistics using point processes.

problem Scarcity of extreme data leads to model misspecification error in EVT.
method Developed DRO estimators informed by semi-parametric max-stable constraints in the space of point processes.
result Proposed DRO estimators improve out-of-sample performance and are validated on synthetic and real data.

New GoF test improves change point detection in multivariate time series.

problem Detecting changes in multivariate time series data efficiently and robustly.
method Developed a novel multivariate rank-energy GoF test (sRE) for change point detection.
result sRE-based CPD outperforms existing methods in AUC and F1-score.

Study on kernel tests for high-dimensional data, focusing on MMD and CLT.

problem Asymptotic behavior of kernel two-sample tests in high dimensions and large samples.
method Maximum mean discrepancy (MMD) with isotropic kernels, deriving asymptotic expansions and CLT.
result Interplay between moment discrepancy and dimension-and-sample orders in kernel tests.