Chiseling finds valid subgroups interactively, improving on existing methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper examines LLM capability benchmarks through construct validity, favoring nomological account.
The article introduces inferential moments for analyzing uncertain multivariable systems.
Post-ADC inference corrects bias in statistical inference after active data collection.
Paper develops methods for statistical inference with SGD in nonconvex optimization.
Prediction of future observations is an important and challenging problem. The two mainstream approaches for quantifying prediction uncertainty use prediction regions and predictive distributions, respectively, with the latter believed to be more informative because it can perform other prediction-related tasks. The st…
Due to the increasing availability of high-dimensional empirical applications in many research disciplines, valid simultaneous inference becomes more and more important. For instance, high-dimensional settings might arise in economic studies due to very rich data sets with many potential covariates or in the analysis o…
Kernel ridge regression inference for nonstandard data.
The paper improves Lasso inference methods for survey data.
Ordinary least square (OLS) estimation of a linear regression model is well-known to be highly sensitive to outliers. It is common practice to (1) identify and remove outliers by looking at the data and (2) to fit OLS and form confidence intervals and p-values on the remaining data as if this were the original data col…
PAIR-CI calibrates CI tests for causal discovery with incomplete data.
New categorization of community detection methods to avoid pitfalls.
This paper proposes a unified framework to quantify local and global inferential uncertainty for high dimensional nonparanormal graphical models. In particular, we consider the problems of testing the presence of a single edge and constructing a uniform confidence subgraph. Due to the presence of unknown marginal trans…
We propose strategies to estimate and make inference on key features of heterogeneous effects in randomized experiments. These key features include best linear predictors of the effects using machine learning proxies, average effects sorted by impact groups, and average characteristics of most and least impacted units.…
We consider the problem of undirected graphical model inference. In many applications, instead of perfectly recovering the unknown graph structure, a more realistic goal is to infer some graph invariants (e.g., the maximum degree, the number of connected subgraphs, the number of isolated nodes). In this paper, we propo…
With the proliferation of mobile devices and the internet of things, developing principled solutions for privacy in time series applications has become increasingly important. While differential privacy is the gold standard for database privacy, many time series applications require a different kind of guarantee, and a…
Study trade-offs between statistical and computational efficiency in variational inference.
Paper introduces ML for rare-event prediction in patent quality estimation.
Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, and Newey (2016) provide a generic double/de-biased machine learning (DML) approach for obtaining valid inferential statements about focal parameters, using Neyman-orthogonal scores and cross-fitting, in settings where nuisance parameters are estimated using a new gene…
This paper explores using SSIM for better image generation in generative models.
Bayesian uncertainty quantification is flawed, according to new research.
This paper develops embeddings that preserve likelihood-based statistical inference.
Gradient-flow optimization is reinterpreted as a statistical inference problem.
This paper concerns the development of an inferential framework for high-dimensional linear mixed effect models. These are suitable models, for instance, when we have repeated measurements for subjects. We consider a scenario where the number of fixed effects is large (and may be larger than ), but the n…
Framework assesses variable importance for heterogeneous treatment effects.
New methods improve inference after prediction without strong model assumptions.
The paper proposes a method to test properties of the optimal assortment in multinomial logit models.
Paper develops conformalized survival analysis method for better prediction.
Using first principles from inference, we design a set of functionals for the purposes of \textit{ranking} joint probability distributions with respect to their correlations. Starting with a general functional, we impose its desired behaviour through the \textit{Principle of Constant Correlations} (PCC), which constrai…
New method for debiased inference without assuming exact solutions in inverse problems.
Bayesian reinforcement learning (BRL) offers a decision-theoretic solution for reinforcement learning. While "model-based" BRL algorithms have focused either on maintaining a posterior distribution on models or value functions and combining this with approximate dynamic programming or tree search, previous Bayesian "mo…
DRF improves confidence and uncertainty assessment for multivariate conditional distributions.
Measures dependence between two systems using Bayesian model comparison.
Paper introduces a new robust method for estimating Pareto tail index from grouped data.
Deep learning methods continue to have a decided impact on machine learning, both in theory and in practice. Statistical theoretical developments have been mostly concerned with approximability or rates of estimation when recovering infinite dimensional objects (curves or densities). Despite the impressive array of ava…
DARTS optimizes covariate selection in trials with limited data.
Understanding the effect of a particular treatment or a policy pertains to many areas of interest, ranging from political economics, marketing to healthcare. In this paper, we develop a non-parametric algorithm for detecting the effects of treatment over time in the context of Synthetic Controls. The method builds on c…
We develop a new modeling framework for Inter-Subject Analysis (ISA). The goal of ISA is to explore the dependency structure between different subjects with the intra-subject dependency as nuisance. It has important applications in neuroscience to explore the functional connectivity between brain regions under natural …
The paper explores how invertibility affects the complexity of encoder models in VAEs.
iWGAN improves GANs by stabilizing training and preventing mode collapse.
We propose a robust inferential procedure for assessing uncertainties of parameter estimation in high-dimensional linear models, where the dimension can grow exponentially fast with the sample size . Our method combines the de-biasing technique with the composite quantile function to construct an estimator that …
Breiman's paper sparked debate on the future of statistics and machine learning.
We study Granger causality testing for high-dimensional time series using regularized regressions. To perform proper inference, we rely on heteroskedasticity and autocorrelation consistent (HAC) estimation of the asymptotic variance and develop the inferential theory in the high-dimensional setting. To recognize the ti…
We propose a new inferential framework for constructing confidence regions and testing hypotheses in statistical models specified by a system of high dimensional estimating equations. We construct an influence function by projecting the fitted estimating equations to a sparse direction obtained by solving a large-scale…
Study uncovers statistical optimality of nonconvex tensor completion methods.
We derive streamlined mean field variational Bayes algorithms for fitting linear mixed models with crossed random effects. In the most general situation, where the dimensions of the crossed groups are arbitrarily large, streamlining is hindered by lack of sparseness in the underlying least squares system. Because of th…
In the 70s a novel branch of statistics emerged focusing its effort in selecting a function in the pattern recognition problem, which fulfils a definite relationship between the quality of the approximation and its complexity. These data-driven approaches are mainly devoted to problems of estimating dependencies with l…
Imputation-Powered Inference improves subpopulation efficiency in missing data settings.