The article introduces inferential moments for analyzing uncertain multivariable systems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
iWGAN improves GANs by stabilizing training and preventing mode collapse.
New categorization of community detection methods to avoid pitfalls.
This paper proposes a unified framework to quantify local and global inferential uncertainty for high dimensional nonparanormal graphical models. In particular, we consider the problems of testing the presence of a single edge and constructing a uniform confidence subgraph. Due to the presence of unknown marginal trans…
With the proliferation of mobile devices and the internet of things, developing principled solutions for privacy in time series applications has become increasingly important. While differential privacy is the gold standard for database privacy, many time series applications require a different kind of guarantee, and a…
Chiseling finds valid subgroups interactively, improving on existing methods.
Study trade-offs between statistical and computational efficiency in variational inference.
This paper explores using SSIM for better image generation in generative models.
New DR method improves robustness in high-dimensional treatment effects.
Stochastic gradient descent (SGD) is an immensely popular approach for online learning in settings where data arrives in a stream or data sizes are very large. However, despite an ever-increasing volume of work on SGD, much less is known about the statistical inferential properties of SGD-based predictions. Taking a fu…
Paper examines LLM capability benchmarks through construct validity, favoring nomological account.
Using first principles from inference, we design a set of functionals for the purposes of \textit{ranking} joint probability distributions with respect to their correlations. Starting with a general functional, we impose its desired behaviour through the \textit{Principle of Constant Correlations} (PCC), which constrai…
Improves Bayesian predictive performance in misspecified models.
Bayesian reinforcement learning (BRL) offers a decision-theoretic solution for reinforcement learning. While "model-based" BRL algorithms have focused either on maintaining a posterior distribution on models or value functions and combining this with approximate dynamic programming or tree search, previous Bayesian "mo…
The growing size of modern data brings many new challenges to existing statistical inference methodologies and theories, and calls for the development of distributed inferential approaches. This paper studies distributed inference for linear support vector machine (SVM) for the binary classification task. Despite a vas…
The paper explores how invertibility affects the complexity of encoder models in VAEs.
New method compares community detection algorithms without ground truth.
In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violate…
Breiman's paper sparked debate on the future of statistics and machine learning.
This paper defines systematic value investing as an empirical optimization problem. Predictive modeling is introduced as a systematic value investing methodology with dynamic and optimization features. A predictive modeling process is demonstrated using financial metrics from Gray & Carlisle and Buffett & Clark. A 31-y…
We derive streamlined mean field variational Bayes algorithms for fitting linear mixed models with crossed random effects. In the most general situation, where the dimensions of the crossed groups are arbitrarily large, streamlining is hindered by lack of sparseness in the underlying least squares system. Because of th…
Bayesian DL model improves DCMs for better predictive and inferential performance.
New approaches improve uncertainty quantification in autoregressive models for sequence data.
PAIR-CI calibrates CI tests for causal discovery with incomplete data.
Paper establishes statistical inference for performative predictions.
New synthetic data analysis reveals high type 1 error rates.
The paper proposes a method to test properties of the optimal assortment in multinomial logit models.
Ordinary least square (OLS) estimation of a linear regression model is well-known to be highly sensitive to outliers. It is common practice to (1) identify and remove outliers by looking at the data and (2) to fit OLS and form confidence intervals and p-values on the remaining data as if this were the original data col…
UCB algorithm provides stable sample means for sequential data.
Paper introduces ML for rare-event prediction in patent quality estimation.
Paper develops methods for statistical inference with SGD in nonconvex optimization.
New approach to topic modelling with covariates for large text corpora.
We propose a likelihood ratio based inferential framework for high dimensional semiparametric generalized linear models. This framework addresses a variety of challenging problems in high dimensional data analysis, including incomplete data, selection bias, and heterogeneous multitask learning. Our work has three main …
FNNs can be made more interpretable with statistical methods.
We consider the problem of undirected graphical model inference. In many applications, instead of perfectly recovering the unknown graph structure, a more realistic goal is to infer some graph invariants (e.g., the maximum degree, the number of connected subgraphs, the number of isolated nodes). In this paper, we propo…
Adaptive Bayesian learning aggregates experts to improve performance.
We present a general framework for classifying partially observed dynamical systems based on the idea of learning in the model space. In contrast to the existing approaches using model point estimates to represent individual data items, we employ posterior distributions over models, thus taking into account in a princi…
Simple method for estimating missing panel data entries with confidence intervals.
Recent decades have seen an interest in prediction problems for which Bayesian methodology has been used ubiquitously. Sampling from or approximating the posterior predictive distribution in a Bayesian model allows one to make inferential statements about potentially observable random quantities given observed data. Th…
Unified framework for predicting data changes influenced by predictions.
New method improves model explainability.
In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here, we consider supervised-learning settings: predicting a target when missing valu…
How should statistical procedures be designed so as to be scalable computationally to the massive datasets that are increasingly the norm? When coupled with the requirement that an answer to an inferential question be delivered within a certain time budget, this question has significant repercussions for the field of s…
Markov jump processes (MJPs) are used to model a wide range of phenomena from disease progression to RNA path folding. However, maximum likelihood estimation of parametric models leads to degenerate trajectories and inferential performance is poor in nonparametric models. We take a small-variance asymptotics (SVA) appr…
The application of existing methods for constructing optimal dynamic treatment regimes is limited to cases where investigators are interested in optimizing a utility function over a fixed period of time (finite horizon). In this manuscript, we develop an inferential procedure based on temporal difference residuals for …
The Mondrian process represents an elegant and powerful approach for space partition modelling. However, as it restricts the partitions to be axis-aligned, its modelling flexibility is limited. In this work, we propose a self-consistent Binary Space Partitioning (BSP)-Tree process to generalize the Mondrian process. Th…
This work improves mixing rates for Bayesian CART, a key component of BART.
Client appraisal improves efficiency in microfinance banks in Adamawa State.