Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2875758621,149 · Jun 202019922001200920182026
48 results for causal data science

Causal inference is crucial for understanding data in Data Science.

problem Understanding causal effects in data science, even when data is non-causal.
method Review of causal roadmap, including scientific question, causal model, estimands, statistical estimators, and interpretation.
result Using the causal roadmap framework improves statistical analysis and interpretation in Data Science.

BCF models estimate causal effects on multiple outcomes in TIMSS data.

problem Estimating causal effects on multiple outcomes in educational data.
method Bayesian Additive Regression Trees (BART) for multivariate causal inference.
result Positive and negative effects of home study conditions and school absence on student achievement.

This paper analyzes social influence using causal data science.

problem Separating genuine causal processes from spurious correlations in social influence data.
method The approach involves partitioning data into groups with minimal contradiction, followed by constrained MLE for causal topology learning.
result The method can retrieve genuine causal arcs and improve influence spread prediction.

SLdisco uses supervised learning to discover causal models from observational data.

problem Estimating causal effects from observational data with limited samples and sparse models.
method Supervised machine learning to map observational data to causal equivalence classes.
result SLdisco is more conservative, less sensitive to sample size, and provides better model inference.

SCIENCE improves prediction intervals for individual causal effects.

problem Wide prediction intervals limit practical utility of causal inference.
method Surrogate-assisted conformal inference for efficient individual causal effects.
result SCIENCE produces more efficient prediction intervals for individual causal effects.

Develops a machine learning pipeline for learning causal structure in time-series data.

problem Current ML algorithms fail to learn causal structure in time-series data due to lack of temporal order consideration.
method Integrates machine learning with chaos theory using ChaosFEX feature extractor to learn generalized causal structure.
result Successfully learns generalized causal structure in time-series data.

New algorithms predict causal links better than traditional methods in time series data.

problem Learning causal structure from time series data with challenges in real-world Earth sciences.
method Combination of established ideas for linear methods to identify causal links in non-linear systems, with a focus on large regression coefficients.
result Large regression coefficients can predict causal links better than small p-values in practice.

Causal relationships in time series with latent variables are discovered using LPCMCI.

problem Discovering causal relationships in complex, time-series data with hidden variables.
method Evaluated LPCMCI algorithm for finding generators compatible with multi-dimensional, autocorrelated time series with latent variables.
result LPCMCI performs better than random guessing but is not optimal.

Responds to critiques on tests for causal parameter confidence intervals.

problem Testing nominal confidence interval coverage for causal parameters estimated by machine learning.
method Rejoinder to critiques on nearly assumption-free tests.
result Clarifies and supports the original research's approach.

New benchmark tests machine learning's ability to learn causal overhypotheses.

problem Machine learning's difficulty in understanding causal overhypotheses.
method Adapted blicket detector environment for machine learning agents to test causal overhypotheses.
result Many state-of-the-art methods struggle with causal overhypotheses in the new benchmark.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

S-DIDML integrates structural DID with ML for causal inference in high-dimensional data.

problem Causal inference in high-dimensional observational panel data with confounding variables.
method Structural identification with high-dimensional estimation, Neyman orthogonality, cross-fitting, causal forests, semi-parametric models.
result Precision in identifying policy-sensitive groups and optimizing resource allocation.

DeepCausalMMM models marketing impacts using deep learning and causal inference.

problem Traditional MMM approaches struggle with non-linear dynamics and temporal patterns.
method Combines deep learning, causal inference, and marketing science. Uses GRUs for temporal patterns and DAG structure for channel dependencies.
result Captures non-linear dynamics and temporal patterns in marketing impacts.

New method identifies nonstationary causal structures in time series data.

problem Identifying causal relationships in time series data that change over time.
method High-order Markov Switching Models for regime-dependent causal discovery.
result Scalable approach for estimating high-order regime-dependent causal structures.

Machine learning's data-centric philosophy conflicts with natural sciences' standards.

problem Conflict between machine learning's ontology and epistemology and natural sciences' practices.
method Identifying and analyzing contexts where ML can be beneficial or harmful in natural sciences.
result ML can enhance trustworthiness in causal inference but introduces biases in emulation and labeling.

Cluster-DAGs improve causal discovery with prior knowledge.

problem Finding cause-effect relationships from high-dimensional data.
method Cluster-DAGs as prior knowledge framework, modified constraint-based algorithms Cluster-PC and Cluster-FCI.
result Cluster-PC and Cluster-FCI outperform baselines without prior knowledge.

GIT uses gradient estimators to target interventions for causal discovery.

problem Challenges in inferring causal structure from observational data.
method GIT uses gradient estimators to target interventions for causal discovery.
result GIT performs on par with competitive baselines, surpassing them in low-data regimes.

Discovering the causal structure among a set of variables is a fundamental problem in many areas of science. In this paper, we propose Kernel Conditional Deviance for Causal Inference (KCDC) a fully nonparametric causal discovery method based on purely observational data. From a novel interpretation of the notion of as…

2018-04-12abs ↗pdf ↗

Develops tools to decompose spurious variations in causal models.

problem Understanding and decomposing spurious variations in causal relationships.
method Formal tools for decomposing spurious effects in Markovian and Semi-Markovian models.
result First results on non-parametric decomposition of spurious effects and sufficient conditions for identification.

CAnDOIT discovers causal relationships using both observational and interventional time-series data.

problem Identifying causal relationships in the presence of hidden factors.
method CAnDOIT combines observational and interventional time-series data to reconstruct causal models.
result CAnDOIT effectively handles interventional data and enhances the accuracy of causal analysis.

New benchmarks show LLMs struggle with causal discovery.

problem Leveraging LLMs for causal discovery is unreliable due to dataset leakage.
method Developing science-grounded benchmarks and hybrid methods combining LLM predictions with statistical analysis.
result LLMs perform poorly on novel, real-world scientific studies compared to classical methods.

The paper develops methods to bound causal effects using Partial Ancestral Graphs.

problem Bounding causal effects from observational data when true causal diagrams are unknown.
method Proposes a method using Partial Ancestral Graphs to derive bounds on causal effects from observational data.
result Demonstrates the effectiveness of the method with synthetic and real data examples.

Framework isolates causal effects from time series data, improving accuracy under non-stationarity and autocorrelation.

problem Causal inference in non-stationary, autocorrelated time series data.
method Decomposes time series into trend, seasonal, and residual components; performs component-specific causal analysis.
result Framework more accurately recovers ground-truth causal structure than state-of-the-art baselines, especially under strong non-stationarity and temporal autocorrelation.

Tree-Query uses LLMs to discover causal relationships in a transparent, interpretable manner.

problem Error propagation in classical causal discovery methods and opaque, confidence-free behavior of recent LLM-based causal oracles.
method Tree-Query is a tree-structured, multi-expert LLM framework that reduces causal discovery to queries about backdoor paths and dependencies.
result Tree-Query provides interpretable judgments with robustness-aware confidence scores and improves structural metrics over LLM baselines.

Survey of deep causal models for industrial applications.

problem Estimating causal effects using deep learning.
method Deep causal models map covariates to a representation space and use objective functions for unbiased counterfactual data estimation.
result Comprehensive overview of deep causal models with industry applications.