Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

316192122 · Jun 202019922001200920172026
48 results for correlation vs causation

Counterfactual diagnosis improves medical accuracy and safety.

problem Existing diagnostic algorithms struggle with distinguishing correlation from causation.
method Reformulated diagnosis as a counterfactual inference task and derived new counterfactual diagnostic algorithms.
result Counterfactual diagnostic algorithms significantly improve accuracy and safety compared to standard Bayesian algorithms.

Neural Shadow-Mapping uncovers causal links in dynamic systems.

problem Discovering causal structures in dynamic systems with mirage correlations.
method Neural network based method embedding high-dimensional data into a shadow representation for causal link estimation.
result Demonstrates performance in discovering causal links from video-representations of dynamic systems.

AI models forget statistics' lesson: correlation doesn't imply causation.

problem AI models often produce flawed causal models due to ignoring correlation vs causation.
method Demonstrates examples of flawed AI models and proposes rethinking core models.
result Current efforts to make AI models ethical are insufficient.

Framework identifies causal factors of climate change using correlations and machine learning.

problem Understanding socioeconomic factors influencing carbon emissions and climate change.
method Three-step framework: correlation analysis, causal discovery, LLM interpretations.
result Adaptable solutions for data-driven policy-making and strategic decision-making.

Paper simplifies calculating causation probabilities and ranks root causes.

problem Computational challenges in assessing causal relationships.
method Algorithmic simplifications and novel methodological framework for Root Cause Analysis.
result Significantly reduces computational complexity for calculating causation probabilities.

CausalBench aims to advance causal learning research with a transparent platform.

problem Lack of unified benchmark datasets, algorithms, metrics, and evaluation interfaces for causal learning.
method Introduces CausalBench, a flexible benchmark framework for causal analysis and machine learning.
result Promotes scientific collaboration, reproducibility, and awareness in causal learning research.

Proposes Causal Loss to improve machine learning models' causal inference.

problem Machine learning algorithms often fail to capture causal relationships when data is inconsistent.
method Introduces Causal Loss, a model-agnostic loss function that enhances interventional capabilities.
result Causal Loss improves non-causal associative models to have interventional capabilities.

Efficiently constructs sparse ROMs for high-dimensional data using causation entropy.

problem Creating effective reduced-order models for high-dimensional dynamical data.
method Uses causation entropy to identify important terms and construct ROMs with varying sparsity.
result Demonstrates the effectiveness of causation entropy in constructing sparse ROMs for chaotic systems with skewed statistics.

Models predict probabilities of causation from limited data.

problem Estimating probabilities of causation requires unreliable or impractical experimental and observational data.
method Proposed Exact-MLP and Mask-MLP models trained on reliable subpopulations.
result Models achieve average MAEs of roughly 0.03, reducing MAE by 80%.

Paper proposes learning causal graphs with only relevant variables.

problem Discovering causal relationships in large-scale graphs often includes irrelevant variables.
method Developed NSCSL algorithm to learn necessary and sufficient causal graphs (NSCG).
result NSCSL algorithm identifies relevant causal features for specific outcomes.

Many supervised learning tasks are emerged in dual forms, e.g., English-to-French translation vs. French-to-English translation, speech recognition vs. text to speech, and image classification vs. image generation. Two dual tasks have intrinsic connections with each other due to the probabilistic correlation between th…

2017-07-03abs ↗pdf ↗

Mastering the dynamics of social influence requires separating, in a database of information propagation traces, the genuine causal processes from temporal correlation, i.e., homophily and other spurious causes. However, most studies to characterize social influence, and, in general, most data-science analyses focus on…

2018-08-06abs ↗pdf ↗

We extend multi-way, multivariate ANOVA-type analysis to cases where one covariate is the view, with features of each view coming from different, high-dimensional domains. The different views are assumed to be connected by having paired samples; this is a common setup in recent bioinformatics experiments, of which we a…

2009-12-16abs ↗pdf ↗

New method tightens bounds on causation probabilities using independent datasets.

problem Challenging point identification of causation probabilities without strong assumptions.
method Imposes counterfactual consistency between SCMs constructed from independent datasets and uses conditional mutual information.
result Significantly tighter bounds on causation probabilities are established.

Critiques causal reductionism in financial studies, suggesting alternative approaches.

problem Limitations of unidirectional causation in self-referencing systems like finance.
method Critical assessment of causal inference in empirical finance, using ecological models.
result Current financial tools may be limited to ex post inference, especially in reflexive contexts.

Study shows gaps in Bitcoin order book are linked to returns but only in the short term.

problem Understanding the relationship between gaps and returns in Bitcoin order books.
method Examined the dynamics of gaps and returns in a Bitcoin order book without considering long-term causation.
result The causal relationship between gaps and returns is limited to instantaneous causation.

New PEMs improve network inference from time-series data.

problem Causal inference from time-series data with trade-off between accuracy and feasibility.
method Infer networks via process motifs for lagged correlation in linear stochastic processes.
result Proposed PEMs achieve high accuracy and efficiency in network inference.

Calibrates historical and implied correlations in energy markets.

problem Challenges in aligning historical correlations of futures contracts with implied volatility smiles.
method Multiplicative multi-factor Heath-Jarrow-Morton model combined with stochastic volatility from lifted Heston model, using Kemna-Vorst approximation and Fourier-based techniques.
result Remarkable joint historical and implied calibration fits on the German power market.

New pricing framework allocates costs of operating reserves and transmission.

problem Allocating costs of operating reserves and transmission efficiently.
method Causation-based framework using contingency-constrained scheduling models.
result More comprehensive and efficient cost-reflective market operations.

Unified model improves multi-task learning by accounting for temporal misalignment.

problem Poor predictive performance and uncertainty quantification due to temporal misalignment in multi-task learning.
method Uses Gaussian processes to model correlations and includes a monotonic warp of the input data to account for temporal misalignment.
result Improves predictive performance and uncertainty quantification in multi-task learning.

AI needs causal inference to avoid being just a correlation machine.

problem AI's inability to distinguish correlation from causation.
method Develops a unified framework connecting various causal statistical estimators and proves a Statistical Necessity Theorem for causal generalization.
result AI systems without causal grounding are brittle and biased, highlighting the need for causal statistics.

A simple guide to understanding hierarchical causality in complex systems.

problem Understanding hierarchical causality in complex systems.
method Formalizing hierarchical causality in terms of actors and agents, with three key structures.
result The system requires three additional structures: causation classes, aggregation operators, and discrete event-time maps.

We are interested in learning causal relationships between pairs of random variables, purely from observational data. To effectively address this task, the state-of-the-art relies on strong assumptions regarding the mechanisms mapping causes to effects, such as invertibility or the existence of additive noise, which on…

2014-09-15abs ↗pdf ↗

A new approach to rationalization identifies true rationales by considering causal relationships.

problem Existing rationalization methods struggle with spuriousness, where snippets with similar contributions are hard to distinguish.
method The method leverages causal inference to identify non-spurious rationales, defining probabilities of causation based on a structural causal model.
result The proposed causal rationalization outperforms existing methods on real-world datasets.

Novel framework combines tree-based discretization and ILP matching for causal inference.

problem Challenges in identifying causal relationships from observational data.
method Combines tree-based discretization and ILP matching for causal inference.
result Yields computational efficiency and less biased ATT estimates.

Model shows triangular arbitrage key to cross-currency correlations in forex markets.

problem Understanding cross-currency correlations in forex markets.
method Agent-based model of market interactions.
result Triangular arbitrage is primary driver of cross-currency correlations.

Study finds carbon emissions affect stock value, but not bought emissions.

problem Determining if carbon emissions impact stock value and whether this is due to direct or indirect emissions.
method Fixed-effects analysis with propensity score weighting to control for selection bias.
result Firms with higher Scope 1 emissions have a statistically significant positive carbon premium, but Scope 2 emissions do not.

Paper tackles P vs NP problem in portfolio optimization with cardinality constraints and Black-Scholes derivatives.

problem Operationalizing the P vs NP problem in cardinality-constrained portfolio selection.
method Mixed-integer quadratic program with genetic algorithms, Monte Carlo sampling, and greedy screening.
result Cardinality constraint reshapes efficient frontier, highlighting trade-offs between stability and computational cost.

FinCARE combines financial data and AI reasoning to improve causal analysis of financial performance.

problem Correlation-based analysis fails to capture true causal relationships in financial performance.
method Hybrid framework integrating causal discovery algorithms with financial domain knowledge from SEC filings and LLM reasoning.
result KG+LLM-enhanced methods improve causal discovery across PC, GES, and NOTEARS by 36-366%.

A new method for sparse Gaussian process regression using correlated experts.

problem Sparse Gaussian process regression for large datasets with cubic computational complexity.
method Aggregating predictions from correlated experts to improve scalability and accuracy.
result Superior performance compared to state-of-the-art methods for synthetic and real-world datasets.

We study the problem of discovering the simplest latent variable that can make two observed discrete variables conditionally independent. The minimum entropy required for such a latent is known as common entropy in information theory. We extend this notion to Renyi common entropy by minimizing the Renyi entropy of the …

2018-07-26abs ↗pdf ↗

A networked learning method for correlated data outperforms federated learning in precision.

problem Estimating models from correlated data distributed across a network.
method Local linear model estimation with network regularization and information exchange.
result The weighted ensemble average estimate converges faster and more precisely than federated learning.

Study on collaboration vs. independent data collection in sensor networks.

problem Impact of sensor correlation on data collection strategies.
method Analysis of Fisher information and Cramer-Rao bound.
result Optimal strategy involves transferring non-immediate information for improved estimation.

Study shows statistical biases can mislead transformer models, impairing their generalization.

problem Statistical biases in transformers affect their ability to generalize.
method Evaluated transformer models on synthetic algorithmic tasks with varying statistical biases.
result Statistical biases lead to overestimation of transformer models' generalization capabilities.

CausalGame benchmarks LLM agents' causal thinking in games.

problem Evaluating causal thinking in AI Scientists with LLMs.
method Interactive games with 14 scenarios incorporating selection bias, measurement error, and hidden confounders.
result None of the 30 LLM agents demonstrated reliable causal thinking, with the best model achieving only 68.0% survival.

Improved stock selection through predictive fundamentals and uncertainty estimates.

problem Selecting stocks based on future financial data to outperform traditional factor models.
method Train deep nets to forecast future fundamentals, incorporate uncertainty estimates, and adjust portfolios to manage risk.
result Simulated annualized return of 17.7% and Sharpe ratio of 0.84 for uncertainty-aware model, significantly higher than 14.0% and 0.52 for standard factor models.

This study uses causal Shapley values to analyze how socioeconomic factors cause the spread of COVID-19.

problem Understanding how socioeconomic factors cause the spread of COVID-19.
method The study employs an explanatory framework from cooperative game theory augmented with do calculus, specifically causal Shapley values, to analyze the causal connections.
result The causal Shapley values reveal distinct advantages of non-linear machine learning models over linear models in multivariate analysis.

This work proposes a method to learn graph structure for multivariate time series forecasting.

problem Improving multivariate time series forecasting by leveraging pairwise information.
method Learning a probabilistic graph model through optimizing mean performance over graph distribution parameterized by a neural network.
result Our method outperforms existing approaches in simplicity, efficiency, and performance.