The article introduces inferential moments for analyzing uncertain multivariable systems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New categorization of community detection methods to avoid pitfalls.
This paper proposes a unified framework to quantify local and global inferential uncertainty for high dimensional nonparanormal graphical models. In particular, we consider the problems of testing the presence of a single edge and constructing a uniform confidence subgraph. Due to the presence of unknown marginal trans…
With the proliferation of mobile devices and the internet of things, developing principled solutions for privacy in time series applications has become increasingly important. While differential privacy is the gold standard for database privacy, many time series applications require a different kind of guarantee, and a…
Chiseling finds valid subgroups interactively, improving on existing methods.
Study trade-offs between statistical and computational efficiency in variational inference.
This paper explores using SSIM for better image generation in generative models.
Using first principles from inference, we design a set of functionals for the purposes of \textit{ranking} joint probability distributions with respect to their correlations. Starting with a general functional, we impose its desired behaviour through the \textit{Principle of Constant Correlations} (PCC), which constrai…
Bayesian reinforcement learning (BRL) offers a decision-theoretic solution for reinforcement learning. While "model-based" BRL algorithms have focused either on maintaining a posterior distribution on models or value functions and combining this with approximate dynamic programming or tree search, previous Bayesian "mo…
The paper explores how invertibility affects the complexity of encoder models in VAEs.
iWGAN improves GANs by stabilizing training and preventing mode collapse.
Breiman's paper sparked debate on the future of statistics and machine learning.
We derive streamlined mean field variational Bayes algorithms for fitting linear mixed models with crossed random effects. In the most general situation, where the dimensions of the crossed groups are arbitrarily large, streamlining is hindered by lack of sparseness in the underlying least squares system. Because of th…
Paper examines LLM capability benchmarks through construct validity, favoring nomological account.
New synthetic data analysis reveals high type 1 error rates.
New DR method improves robustness in high-dimensional treatment effects.
Ordinary least square (OLS) estimation of a linear regression model is well-known to be highly sensitive to outliers. It is common practice to (1) identify and remove outliers by looking at the data and (2) to fit OLS and form confidence intervals and p-values on the remaining data as if this were the original data col…
UCB algorithm provides stable sample means for sequential data.
Paper develops methods for statistical inference with SGD in nonconvex optimization.
New approach to topic modelling with covariates for large text corpora.
We propose a likelihood ratio based inferential framework for high dimensional semiparametric generalized linear models. This framework addresses a variety of challenging problems in high dimensional data analysis, including incomplete data, selection bias, and heterogeneous multitask learning. Our work has three main …
Stochastic gradient descent (SGD) is an immensely popular approach for online learning in settings where data arrives in a stream or data sizes are very large. However, despite an ever-increasing volume of work on SGD, much less is known about the statistical inferential properties of SGD-based predictions. Taking a fu…
We consider the problem of undirected graphical model inference. In many applications, instead of perfectly recovering the unknown graph structure, a more realistic goal is to infer some graph invariants (e.g., the maximum degree, the number of connected subgraphs, the number of isolated nodes). In this paper, we propo…
Simple method for estimating missing panel data entries with confidence intervals.
Recent decades have seen an interest in prediction problems for which Bayesian methodology has been used ubiquitously. Sampling from or approximating the posterior predictive distribution in a Bayesian model allows one to make inferential statements about potentially observable random quantities given observed data. Th…
The growing size of modern data brings many new challenges to existing statistical inference methodologies and theories, and calls for the development of distributed inferential approaches. This paper studies distributed inference for linear support vector machine (SVM) for the binary classification task. Despite a vas…
New method improves model explainability.
In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here, we consider supervised-learning settings: predicting a target when missing valu…
How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…
In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violate…
This work improves mixing rates for Bayesian CART, a key component of BART.
This paper develops embeddings that preserve likelihood-based statistical inference.
New approaches improve uncertainty quantification in autoregressive models for sequence data.
New tools connect CP to GF inference for better probabilistic prediction.
The paper improves Lasso inference methods for survey data.
Deep learning (DL) is a high dimensional data reduction technique for constructing high-dimensional predictors in input-output models. DL is a form of machine learning that uses hierarchical layers of latent features. In this article, we review the state-of-the-art of deep learning from a modeling and algorithmic persp…
Post-ADC inference corrects bias in statistical inference after active data collection.
Proposes a semi-Bayesian nonparametric estimator for MMD in GOF tests and GANs.
PAIR-CI calibrates CI tests for causal discovery with incomplete data.
Privacy-preserving synthetic data from EHRs for learning and inference.
Develops tools to audit ML models for bias and unfairness.
We propose strategies to estimate and make inference on key features of heterogeneous effects in randomized experiments. These key features include best linear predictors of the effects using machine learning proxies, average effects sorted by impact groups, and average characteristics of most and least impacted units.…
The paper proposes a method to test properties of the optimal assortment in multinomial logit models.
NoFAS combines variational inference and adaptive surrogate models for efficient inference of computationally expensive models.
Profile graphical models represent multivariate dependence under varying risk factors.
Bayesian DL model improves DCMs for better predictive and inferential performance.
Improves Bayesian predictive performance in misspecified models.
We present the Causal Gaussian Process Convolution Model (CGPCM), a doubly nonparametric model for causal, spectrally complex dynamical phenomena. The CGPCM is a generative model in which white noise is passed through a causal, nonparametric-window moving-average filter, a construction that we show to be equivalent to …