Develops large-sample theory for non-stationary source separation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Infinitesimal boosting converges to a deterministic process in large sample limit.
We study the distribution of the adaptive LASSO estimator (Zou (2006)) in finite samples as well as in the large-sample limit. The large-sample distributions are derived both for the case where the adaptive LASSO estimator is tuned to perform conservative model selection as well as for the case where the tuning results…
New method compresses large sample data for faster discriminant analysis.
The paper analyzes LIME for tabular data and proves its behavior in large samples.
We study the distributions of the LASSO, SCAD, and thresholding estimators, in finite samples and in the large-sample limit. The asymptotic distributions are derived for both the case where the estimators are tuned to perform consistent model selection and for the case where the estimators are tuned to perform conserva…
We derive formulas for F measures' standard error and confidence intervals.
New tuning rules for Metropolis algorithms derived from Bayesian large-sample asymptotics.
Large sample size brings the computation bottleneck for modern data analysis. Subsampling is one of efficient strategies to handle this problem. In previous studies, researchers make more fo- cus on subsampling with replacement (SSR) than on subsampling without replacement (SSWR). In this paper we investigate a kind of…
Operational risk models commonly employ maximum likelihood estimation (MLE) to fit loss data to heavy-tailed distributions. Yet several desirable properties of MLE (e.g. asymptotic normality) are generally valid only for large sample-sizes, a situation rarely encountered in operational risk. In this paper, we study how…
Although consistency is a minimum requirement of any estimator, little is known about consistency of the mean partition approach in consensus clustering. This contribution studies the asymptotic behavior of mean partitions. We show that under normal assumptions, the mean partition approach is consistent and asymptotic …
We analyze SGAs for statistical inference via asymptotics, improving tuning methods.
Financial econometrics has become an increasingly popular research field. In this paper we review a few parametric and nonparametric models and methods used in this area. After introducing several widely used continuous-time and discrete-time models, we study in detail dependence structures of discrete samples, includi…
Generative Adversarial Networks (GANs) are a class of generative algorithms that have been shown to produce state-of-the art samples, especially in the domain of image creation. The fundamental principle of GANs is to approximate the unknown distribution of a given data set by optimizing an objective function through a…
Develops scalable methods to assess sensitivity and uncertainty in continuous treatment effects.
New methods for estimating causal effects with limited overlap, using Stable Probability Weighting.
A new method reduces feature screening cost from to .
Motivated by safety-critical applications, test-time attacks on classifiers via adversarial examples has recently received a great deal of attention. However, there is a general lack of understanding on why adversarial examples arise; whether they originate due to inherent properties of data or due to lack of training …
A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample from the original full sample and uses it as a surrogate for subsequent computati…
Using 1-min returns of Bitcoin prices, we investigate statistical properties and multifractality of a Bitcoin time series. We find that the 1-min return distribution is fat-tailed, and kurtosis largely deviates from the Gaussian expectation. Although for large sampling periods, kurtosis is anticipated to approach the G…
We develop a sequential low-complexity inference procedure for Dirichlet process mixtures of Gaussians for online clustering and parameter estimation when the number of clusters are unknown a-priori. We present an easily computable, closed form parametric expression for the conditional likelihood, in which hyperparamet…
This paper rigorously establishes that the existence of the maximum likelihood estimate (MLE) in high-dimensional logistic regression models with Gaussian covariates undergoes a sharp `phase transition'. We introduce an explicit boundary curve , parameterized by two scalars measuring the overall magnitu…
The subtle and unique imprint of dark matter substructure on extended arcs in strong lensing systems contains a wealth of information about the properties and distribution of dark matter on small scales and, consequently, about the underlying particle physics. However, teasing out this effect poses a significant challe…
Spectral risk measures (SRMs) belong to the family of coherent risk measures. A natural estimator for the class of SRMs has the form of L-statistics. Various authors have studied and derived the asymptotic properties of the empirical estimator of SRM. We propose a kernel based estimator of SRM. We investigate the large…
Unified framework for estimating high-dimensional conditional factor models.
Neural causal discovery methods fail to accurately uncover causal structures due to the faithfulness property.
A practical algorithm improves approximate OT distances using quantization.
This paper introduces online algorithms to estimate robust geometric median in large data streams.
The paper analyzes the excess risk of PCA and provides a precise characterization.
Understanding the pathways whereby an intervention has an effect on an outcome is a common scientific goal. A rich body of literature provides various decompositions of the total intervention effect into pathway specific effects. Interventional direct and indirect effects provide one such decomposition. Existing estima…
DRIVE improves IV estimation by accounting for distributional uncertainties.
The paper analyzes SGD with dropout regularization in linear models, proving asymptotic properties and providing inference tools.
Maximum Variance Unfolding is one of the main methods for (nonlinear) dimensionality reduction. We study its large sample limit, providing specific rates of convergence under standard assumptions. We find that it is consistent when the underlying submanifold is isometric to a convex subset, and we provide some simple e…
Paper offers a simple CDS approximation formula with high accuracy.
New framework analyzes SGD dynamics in large samples and dimensions.
Online (also called "recursive" or "adaptive") estimation of fixed model parameters in hidden Markov models is a topic of much interest in times series modelling. In this work, we propose an online parameter estimation algorithm that combines two key ideas. The first one, which is deeply rooted in the Expectation-Maxim…
This project was motivated by a dialysis study in northern Taiwan. Dialysis patients, after shunt implantation, may experience two types ("acute" or "non-acute") of shunt thrombosis, both of which may recur. We formulate the problem under the framework of recurrent events data in the presence of competing risks. In par…
Estimates causal contributions of multiple causes on outcome changes.
Develops a method for causal inference with noisy confounders.
We present a new package in R implementing Bayesian additive regression trees (BART). The package introduces many new features for data analysis using BART such as variable selection, interaction detection, model diagnostic plots, incorporation of missing data and the ability to save trees for future prediction. It is …
Nyström KPCA balances computational efficiency and statistical accuracy.
New auditors assess -DP privacy with adaptive sampling, avoiding large sample sizes.
Study estimates heterogeneous principal causal effects with binary treatments and intermediate variables.
Understanding the relationships between different properties of data, such as whether a connectome or genome has information about disease status, is becoming increasingly important in modern biological datasets. While existing approaches can test whether two properties are related, they often require unfeasibly large …
Causal effect estimation from observational data is an important and much studied research topic. The instrumental variable (IV) and local causal discovery (LCD) patterns are canonical examples of settings where a closed-form expression exists for the causal effect of one variable on another, given the presence of a th…
New methods for handling confounding in observational studies.
For a finite function class we describe the large sample limit of the sequential Rademacher complexity in terms of the viscosity solution of a -heat equation. In the language of Peng's sublinear expectation theory, the same quantity equals to the expected value of the largest order statistics of a multidimensional $…
Clustering with fast algorithms large samples of high dimensional data is an important challenge in computational statistics. Borrowing ideas from MacQueen (1967) who introduced a sequential version of the -means algorithm, a new class of recursive stochastic gradient algorithms designed for the -medians loss cri…