The article introduces inferential moments for analyzing uncertain multivariable systems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
As datasets capturing human choices grow in richness and scale -- particularly in online domains -- there is an increasing need for choice models that escape traditional choice-theoretic axioms such as regularity, stochastic transitivity, and Luce's choice axiom. In this work we introduce the Pairwise Choice Markov Cha…
New categorization of community detection methods to avoid pitfalls.
This paper proposes a unified framework to quantify local and global inferential uncertainty for high dimensional nonparanormal graphical models. In particular, we consider the problems of testing the presence of a single edge and constructing a uniform confidence subgraph. Due to the presence of unknown marginal trans…
With the proliferation of mobile devices and the internet of things, developing principled solutions for privacy in time series applications has become increasingly important. While differential privacy is the gold standard for database privacy, many time series applications require a different kind of guarantee, and a…
Chiseling finds valid subgroups interactively, improving on existing methods.
Understanding the effect of a particular treatment or a policy pertains to many areas of interest, ranging from political economics, marketing to healthcare. In this paper, we develop a non-parametric algorithm for detecting the effects of treatment over time in the context of Synthetic Controls. The method builds on c…
Study trade-offs between statistical and computational efficiency in variational inference.
Novel framework for Bayesian reinforcement learning infers value function distributions.
New shape representation for airfoils improves design and manufacturing.
This paper explores using SSIM for better image generation in generative models.
Using first principles from inference, we design a set of functionals for the purposes of \textit{ranking} joint probability distributions with respect to their correlations. Starting with a general functional, we impose its desired behaviour through the \textit{Principle of Constant Correlations} (PCC), which constrai…
Inference over tails is usually performed by fitting an appropriate limiting distribution over observations that exceed a fixed threshold. However, the choice of such threshold is critical and can affect the inferential results. Extreme value mixture models have been defined to estimate the threshold using the full dat…
New measures for prediction validity and consonant plausibility introduced.
The paper explores how invertibility affects the complexity of encoder models in VAEs.
iWGAN improves GANs by stabilizing training and preventing mode collapse.
Breiman's paper sparked debate on the future of statistics and machine learning.
We derive streamlined mean field variational Bayes algorithms for fitting linear mixed models with crossed random effects. In the most general situation, where the dimensions of the crossed groups are arbitrarily large, streamlining is hindered by lack of sparseness in the underlying least squares system. Because of th…
The Unlearning algorithm improves neural network performance in memory tasks.
New strategy debiases synthetic data generated by DGMs for improved statistical inference.
New method optimizes treatment policies to avoid winner's curse.
Paper examines LLM capability benchmarks through construct validity, favoring nomological account.
New synthetic data analysis reveals high type 1 error rates.
New DR method improves robustness in high-dimensional treatment effects.
Ordinary least square (OLS) estimation of a linear regression model is well-known to be highly sensitive to outliers. It is common practice to (1) identify and remove outliers by looking at the data and (2) to fit OLS and form confidence intervals and p-values on the remaining data as if this were the original data col…
UCB algorithm provides stable sample means for sequential data.
Paper develops methods for statistical inference with SGD in nonconvex optimization.
New approach to topic modelling with covariates for large text corpora.
We propose a likelihood ratio based inferential framework for high dimensional semiparametric generalized linear models. This framework addresses a variety of challenging problems in high dimensional data analysis, including incomplete data, selection bias, and heterogeneous multitask learning. Our work has three main …
Stochastic gradient descent (SGD) is an immensely popular approach for online learning in settings where data arrives in a stream or data sizes are very large. However, despite an ever-increasing volume of work on SGD, much less is known about the statistical inferential properties of SGD-based predictions. Taking a fu…
We consider the problem of undirected graphical model inference. In many applications, instead of perfectly recovering the unknown graph structure, a more realistic goal is to infer some graph invariants (e.g., the maximum degree, the number of connected subgraphs, the number of isolated nodes). In this paper, we propo…
Simple method for estimating missing panel data entries with confidence intervals.
Recent decades have seen an interest in prediction problems for which Bayesian methodology has been used ubiquitously. Sampling from or approximating the posterior predictive distribution in a Bayesian model allows one to make inferential statements about potentially observable random quantities given observed data. Th…
The growing size of modern data brings many new challenges to existing statistical inference methodologies and theories, and calls for the development of distributed inferential approaches. This paper studies distributed inference for linear support vector machine (SVM) for the binary classification task. Despite a vas…
New method improves model explainability.
In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here, we consider supervised-learning settings: predicting a target when missing valu…
How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…
In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violate…
This work improves mixing rates for Bayesian CART, a key component of BART.
This paper develops embeddings that preserve likelihood-based statistical inference.
New approaches improve uncertainty quantification in autoregressive models for sequence data.
New tools connect CP to GF inference for better probabilistic prediction.
The paper improves Lasso inference methods for survey data.
Deep learning (DL) is a high dimensional data reduction technique for constructing high-dimensional predictors in input-output models. DL is a form of machine learning that uses hierarchical layers of latent features. In this article, we review the state-of-the-art of deep learning from a modeling and algorithmic persp…
Post-ADC inference corrects bias in statistical inference after active data collection.
Proposes a semi-Bayesian nonparametric estimator for MMD in GOF tests and GANs.
PAIR-CI calibrates CI tests for causal discovery with incomplete data.
Privacy-preserving synthetic data from EHRs for learning and inference.