Causal discovery predicts unobserved joint statistics from observed data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper identifies unobserved variables from observable data.
A new method uses randomized trials to estimate the strength of unobserved confounding.
The edge structure of the graph defining an undirected graphical model describes precisely the structure of dependence between the variables in the graph. In many applications, the dependence structure is unknown and it is desirable to learn it from data, often because it is a preliminary step to be able to ascertain c…
KRCD detects unobserved confounders in nonlinear observational data.
Raising statistical hurdles may not be justified due to data bias.
New method scores DAGs by identifying unobserved confounding.
New model improves multimodal autoencoders by learning joint and conditional distributions.
Paper tackles unobserved confounding in human-AI collaborations.
New method estimates policy performance under unobserved confounding.
A probabilistic query may not be estimable from observed data corrupted by missing values if the data are not missing at random (MAR). It is therefore of theoretical interest and practical importance to determine in principle whether a probabilistic query is estimable from missing data or not when the data are not MAR.…
This work highlights problems with off-policy estimation in recommender systems due to unobserved confounders.
We introduce a variant of the Barndorff-Nielsen and Shephard stochastic volatility model where the non Gaussian Ornstein-Uhlenbeck process describes some measure of trading intensity like trading volume or number of trades instead of unobservable instantaneous variance. We develop an explicit estimator based on marting…
We describe a method that infers whether statistical dependences between two observed variables X and Y are due to a "direct" causal link or only due to a connecting causal path that contains an unobserved variable of low complexity, e.g., a binary variable. This problem is motivated by statistical genetics. Given a ge…
We consider the task of learning a parametric Continuous Time Markov Chain (CTMC) sequence model without examples of sequences, where the training data consists entirely of aggregate steady-state statistics. Making the problem harder, we assume that the states we wish to predict are unobserved in the training data. Spe…
Credit risk analysis improved with a joint model for spatial and temporal effects.
New method combines score lists using joint CDFs, improving computation.
Paper learns Cartesian product graphs with Laplacian constraints.
Aggregate network properties such as cluster cohesion and the number of bridge nodes can be used to glean insights about a network's community structure, spread of influence and the resilience of the network to faults. Efficiently computing network properties when the network is fully observed has received significant …
New method for robust policy evaluation in offline reinforcement learning with sequentially exogenous unobserved confounders.
GUM tackles MARL by avoiding overestimation through state-marginal restriction.
Framework improves CATE estimation by aligning active learning with causal objectives.
New method predicts dynamic relationships in terrorist networks.
We provide a distribution-free test that can be used to determine whether any two joint distributions and are statistically different by inspection of a large enough set of samples. Following recent efforts from Long et al. [1], we rely on joint kernel distribution embedding to extend the kernel two-sample test…
Using AI predictions as data can mislead inference, study shows.
Paper uses non-Euclidean analysis to classify brain structure variations.
Monte Carlo Tree Search (MCTS) algorithms have achieved great success on many challenging benchmarks (e.g., Computer Go). However, they generally require a large number of rollouts, making their applications costly. Furthermore, it is also extremely challenging to parallelize MCTS due to its inherent sequential nature:…
Predictive models can fail to generalize from training to deployment environments because of dataset shift, posing a threat to model reliability and the safety of downstream decisions made in practice. Instead of using samples from the target distribution to reactively correct dataset shift, we use graphical knowledge …
Surveying joint Gaussian graphical models to identify shared structures across domains.
Unified framework infers time-varying graphs from incomplete signals.
A statistical test of independence may be constructed using the Hilbert-Schmidt Independence Criterion (HSIC) as a test statistic. The HSIC is defined as the distance between the embedding of the joint distribution, and the embedding of the product of the marginals, in a Reproducing Kernel Hilbert Space (RKHS). It has …
We take a new look at the problem of disentangling the volatility and jumps processes of daily stock returns. We first provide a computational framework for the univariate stochastic volatility model with Poisson-driven jumps that offers a competitive inference alternative to the existing tools. This methodology is the…
AI task delegation faces incentive collapse with unbounded payments as AI accuracy rises.
A two-step nonparametric method estimates financial systemic risk.
Valid causal inference with unobserved confounding in high-dimensional settings.
This paper addresses the problem of identifying a lower dimensional space where observed data can be sparsely represented. This under-complete dictionary learning task can be formulated as a blind separation problem of sparse sources linearly mixed with an unknown orthogonal mixing matrix. This issue is formulated in a…
The paper explores the relationship between joint mixability and negative dependence structures.
Extends SW and GSW to compare heterogeneous joint distributions.
We study the problem of identifying the causal relationship between two discrete random variables from observational data. We recently proposed a novel framework called entropic causality that works in a very general functional model but makes the assumption that the unobserved exogenous variable has small entropy in t…
New method estimates treatment effects over time with unobserved confounders.
Proposes a novel tensor-based approach for multi-level link prediction.
Biased sampling and missing data complicates statistical problems ranging from causal inference to reinforcement learning. We often correct for biased sampling of summary statistics with matching methods and importance weighting. In this paper, we study nearest neighbor matching (NNM), which makes estimates of populati…
New research shows imputation and regression together can predict better than separate steps.
Individual risk models need to capture possible correlations as failing to do so typically results in an underestimation of extreme quantiles of the aggregate loss. Such dependence modelling is particularly important for managing credit risk, for instance, where joint defaults are a major cause of concern. Often, the d…
New method recovers predictions from unobservable source subpopulation in binary classification.
We propose a tensor-based model that fuses a more granular representation of user preferences with the ability to take additional side information into account. The model relies on the concept of ordinal nature of utility, which better corresponds to actual user perception. In addition to that, unlike the majority of h…
CDVAE estimates treatment effects over time by accounting for unobserved variables.
Paper develops a theory explaining contrastive pre-training for multimodal AI.