Transformer-based method for causal discovery with prior knowledge integration.
problem Complex nonlinear dependencies and spurious correlations in time series data.
method Multi-layer Transformer forecaster with gradient-based causal structure extraction and attention masking for prior knowledge integration.
result Significant improvement in causal discovery and causal lag estimation compared to state-of-the-art methods.
Transformers learn causal structure through gradient descent on self-attention mechanisms.
problem Understanding how transformers learn causal structure during training.
method In-context learning task and simplified two-layer transformer model.
result Gradient descent on a simplified transformer learns to encode latent causal graphs.
This work tackles OOD generalization by leveraging causal invariance without needing to recover causal features.
problem Learning models that perform well on out-of-distribution (OOD) data.
method Causal invariant transformations to modify non-causal features while preserving causal parts.
result Theoretical and practical methods to learn a minimax optimal model across domains using single domain data.
Transformer model handles causal inference with DAG integration.
problem Complex causal structures and adaptability across various scenarios.
method Integrates DAGs into transformer's attention mechanism.
result Surpasses existing methods in estimating causal effects.
CaTs use DAGs with transformers to enforce causal constraints, improving neural network robustness.
problem Neural networks lack inherent causal structure respect, leading to reliability issues.
method Introducing Causal Transformers (CaTs) that operate under predefined causal constraints specified by DAGs.
result CaTs improve robustness and interpretability of neural networks under causal constraints.
This paper tackles causal representation learning with linear and general transformations.
problem Identify and recover latent causal variables and graphs under unknown transformations.
method Score-based algorithms that use gradients of log-density functions for identifiability and achievability.
result Two stochastic hard interventions per node are sufficient for identifiability of general transformations.
Paper recovers latent causal structure and linear transformation from indirect observations.
problem Recovering latent causal structure and linear transformation from indirect observations.
method Established sufficient conditions for DAG recovery, leveraged score function properties, and used soft/hard interventions.
result Perfect recovery of latent DAG structure and linear transformation up to scaling using soft interventions, hard interventions with additional hypothesis testing.
Transformer-based method improves causal discovery from observational data.
problem Causal discovery from observational data requires explicit assumptions.
method CSIvA transformer architecture trained on synthetic data.
result Transformer-based methods adhere to identifiability theory.
Based on the recent work \cite{PII} we put forward a new type of transformation for Lorentzian manifolds characterized by mapping every causal future-directed vector onto a causal future-directed vector. The set of all such transformations, which we call causal symmetries, has the structure of a submonoid which contain…
Develops a Causal Transformer for estimating counterfactual outcomes from longitudinal data.
problem Estimating counterfactual outcomes over time from observational data is challenging due to complex, long-range dependencies.
method Combines three transformer subnetworks with separate inputs for time-varying covariates, previous treatments, and previous outcomes into a joint network with in-between cross-attentions. Uses a custom, end-to-end training procedure with a counterfactual domain confusion loss to address confounding bias.
result Achieves superior performance over current baselines in synthetic and real-world datasets.
We define a new type of transformation for Lorentzian manifolds characterized by mapping every causal future-directed vector onto a causal future-directed vector. The set of all such transformations, which we call causal symmetries, has the structure of a submonoid. Some of their properties are investigated and we give…
CSHT predicts financial returns from news using a novel transformer model on a sphere.
problem Financial forecasting from news and sentiment.
method Granger-causal hypergraph structure, Riemannian geometry, causally masked Transformer attention.
result CSHT outperforms baselines in return prediction, regime classification, and asset ranking.
Flow models recover causal transformations from observational data and a valid ordering.
problem Causal inference with only observational data and a valid causal ordering.
method Flow models that can recover component-wise, invertible transformations of exogenous variables.
result Flow models outperform previous methods and deliver consistent performance across various structural causal models.
The study explores how Transformers predict the next token in a sequence.
problem Understanding the mechanism behind Transformers' autoregressive learning ability.
method Exploring the approximation ability of Transformers for next-token prediction through specific instances and a causal kernel descent method.
result Transformer models can learn context-dependent functions f for next-token prediction based on past and current observations. Introduces Causal Energy Minimization to understand Transformer layers.
problem Empirical parameterization of Transformer blocks remains largely unexplored.
method Causal Energy Minimization framework that recasts Transformer layers as optimization steps on conditional energy functions.
result Identifies design space for Transformer layers including weight sharing and energy-based interpretations.
This paper introduces a new task to better understand Transformers in quantitative contexts.
problem Understanding Transformers in high-stakes quantitative and scientific applications.
method Introduces a novel contextual counting task and analyzes it with causal and non-causal Transformer architectures.
result Causal attention is better suited for the contextual counting task, and no positional embeddings lead to the best accuracy.
The paper develops a framework for abstracting causal models using category theory.
problem Difficulties in changing the variables used to describe a system, especially from fine-grained to coarse-grained.
method Introduces a category of interventional causal models and uses enriched category theory to prove compositionality properties.
result Compositionality of model transformations is established, with bounded errors for each step.
Although nonstationary data are more common in the real world, most existing causal discovery methods do not take nonstationarity into consideration. In this letter, we propose a kernel embedding-based approach, ENCI, for nonstationary causal model inference where data are collected from multiple domains with varying d…
New framework for AI to learn causal models through experience.
problem Lack of guidance for variable choice and interventions in causal models for AI.
method Defines actions as state space transformations, introduces causal variables, and identifies interventions.
result Clarifies the concept of interventions and makes causal representation learning clearer.
MOCA uses modular attention to estimate causal effects from complex data.
problem Estimating causal effects from observational data with complex, non-linear, and high-dimensional treatment and outcome mechanisms.
method MOCA is a transformer-based framework that separates treatment and outcome modeling through modular design and one-way attention mechanism, with cutting-feedback to prevent outcome influence on treatment representations.
result MOCA outperforms classical estimators and machine learning approaches across various simulated and real-world scenarios.
We tackle causal inference under conditional moment restrictions using importance weighting.
problem Challenges in causal inference under conditional moment restrictions, especially in high-dimensional settings.
method Transform conditional moment restrictions to unconditional moment restrictions through importance weighting.
result Successfully estimate nonparametric functions defined under conditional moment restrictions.
Causal spacetimes with Ricci tensor have unique transformations.
problem Understanding transformations in viable causal spacetimes.
method Analyzing Ricci tensor and causal diffeomorphisms.
result Causal diffeomorphisms preserving Ricci tensor are homotheties.
The paper tackles causal disentanglement with linear models and interventions.
problem Identify latent variables in a causal model from observed data.
method Use linear transformations and interventions to uniquely identify latent variables.
result A single intervention on each latent variable is sufficient for identifying the latent causal model.
New method uses kernel deviance measures to discover causal relationships in heterogeneous data.
problem Discovering causal relationships in complex, heterogeneous datasets.
method KIIM-HT, a novel score measure based on heterogeneous transformations of RKHS embeddings.
result KIIM-HT outperforms previous methods in causal discovery tasks.
Framework calculates positional influence in causal residual Transformers.
problem Understanding positional influence in causal residual Transformers.
method Adjoint-sensitivity framework for positional influence in causal residual Transformers.
result Exact evolution of adjoint-energy influence density and decomposition into residual transmission, nonlocal Volterra, and local channels.
Synthetic approach to conformal transformations in metric and Lorentzian spaces.
problem Defining consistent conformal transformations in spaces of low regularity.
method Introducing conformal transformations in metric and Lorentzian spaces, focusing on Lorentzian pre-length spaces.
result Established a consistent notion of conformal length and proved its properties.
DAG-FM discovers causal relationships from heterogeneous data.
problem Challenges in causal discovery from heterogeneous causal mechanisms.
method DAG-FM uses two specialized Transformer-based sub-modules and a robust tabular interaction block to model complex row-column interactions.
result DAG-FM achieves state-of-the-art performance on synthetic and real-world datasets.
Timer-XL predicts multidimensional time series using a unified Transformer approach.
problem Unified time series forecasting across various tasks and contexts.
method Decoder-only Transformers with a universal TimeAttention mechanism and deft position embedding.
result State-of-the-art performance across multiple forecasting benchmarks.
GO-CBED optimizes experiments for specific causal queries, improving efficiency.
problem Efficiently infer causal relationships with limited resources.
method Goal-oriented Bayesian framework that maximizes expected information gain on user-specified causal quantities.
result GO-CBED outperforms existing methods in various causal tasks, especially with limited budgets.
CInA method uses attention to improve causal inference.
problem Challenges in causal inference, especially in complex tasks.
method CInA method utilizes self-supervised causal learning with multiple unlabeled datasets and transformer-type architecture.
result CInA effectively generalizes to out-of-distribution datasets and various real-world datasets.
In this work we define and study the relations between Lorentzian Manifolds given by the diffeomorphisms which map causal future directed vectors onto causal future directed vectors. This class of diffeomorphisms, called proper causal relations, contains as a subset the well-known group of conformal relations and are d…
New method learns causal models from data efficiently.
problem Learning Structural Causal Models from data is challenging.
method Amortized inference via Conditional Fixed-Point Iterations with transformer embeddings.
result Single model predicts causal mechanisms conditioned on data and graph.
Causal deep learning tackles causal inference using tensor factor analysis.
problem Addressing causal questions in data using neural networks.
method Tensor factor analysis and neural network architectures (causal capsules, tensor transformer, multilinear projection algorithm).
result Derives deep neural networks for causal inference with tensor factor analysis.
CausalPFN automates causal effect estimation from observational data.
problem Manual selection of causal effect estimators is time-consuming and requires domain expertise.
method CausalPFN is a transformer that learns to infer causal effects from raw observations without task-specific adjustments.
result CausalPFN achieves superior performance on various benchmarks and real-world tasks.
Paper establishes identifiability and achievability for causal representation learning.
problem Identifying and recovering latent causal models and variables from observational and interventional data.
method Establishes identifiability and achievability using uncoupled interventions and a recovery algorithm.
result Guaranteed perfect recovery of latent causal model and variables under uncoupled interventions.
Researchers use DT to transfer policies from one environment to another using causal reasoning.
problem Adapting to changes in environmental dynamics in reinforcement learning.
method Applying causal counterfactual reasoning to Decision Transformer (DT) architecture for policy transfer.
result DT successfully transfers a learned policy to new environments while retaining most of the reward.
New methods for Z-transform inversion and Wiener-Hopf factorization.
problem Efficient numerical inversion of Z-transforms and factorization of functions. method Sinh-deformations of contours, variable changes, and simplified trapezoid rule.
result High precision and speed in evaluating moments and constructing filters.
New framework infers causal shifts in event sequences under out-of-domain interventions.
problem Inferring causal relationships in event sequences without considering out-of-domain interventions.
method Proposes a new causal framework to define ATE, designs an unbiased ATE estimator, and uses a Transformer-based neural network model.
result Demonstrates superior performance in ATE estimation and goodness-of-fit under out-of-domain-augmented point processes.
New model learns causal world dynamics from state space models.
problem Lack of causal world models in neural world modeling.
method State Space Models (SSM) with attention mechanisms.
result SSM can learn causal models of environments with equivalent performance.
We fully develop the concept of causal symmetry introduced in Class. Quant. Grav. 20 (2003) L139. A causal symmetry is a transformation of a Lorentzian manifold (V,g) which maps every future-directed vector onto a future-directed vector. We prove that the set of all causal symmetries is not a group under the usual comp…
New method identifies latent causal factors from observational data alone.
problem Identifying latent causal factors without interventions or graphical restrictions.
method Characterization of latent factors in nonlinear causal models with additive Gaussian noise and linear mixing, using a practical algorithm based on solving a quadratic program over observed data.
result Latent causal variables can be identified up to a layer-wise transformation, and further disentanglement is not possible.
Paper proposes a new method for demand forecasting in pricing contexts.
problem Demand forecasting in pricing contexts, especially in a profit optimal manner.
method Combines Double Machine Learning for causal inference and transformer-based forecasting models.
result Our method outperforms other forecasting methods in off-policy settings.
TRAM-DAG models bridge interpretability and flexibility in causal modeling.
problem Modeling causal relationships in diverse data types while maintaining interpretability.
method Using transformation models (TRAMs) within structural causal models (SCMs) to handle various data types and maintain interpretability.
result TRAM-DAG models achieve equal or superior performance in causal queries across different levels of the causal hierarchy.
A new method infers causal structures and generates data without DAGs.
problem Modeling causal relationships without DAGs.
method Fixed-point approach on causally ordered variables, amortized TO inference, transformer-based SCM learning.
result The model learns TOs and SCMs from data, outperforming baselines.
Framework quantifies financial NLP robustness under regime shifts.
problem Semantic and causal drift in financial news narratives.
method Four metrics: FCAS, PCS, TSV, NLICS.
result Transformer models are more affected by semantic drift.
LANCA uses ANM to learn latent causal factors without supervision.
problem Learning latent causal factors without supervision.
method LANCA employs a deterministic Wasserstein Auto-Encoder coupled with a differentiable ANM Layer.
result LANCA outperforms baselines on physics and photorealistic environments.
New method selects causal features from diverse data types.
problem Discovering causal relationships from non-continuous data types.
method Transformation-Model (TRAM) based Invariant Causal Prediction (TRAM-ICP) with TRAM-GCM and TRAM-Wald tests.
result Improved power and type I error control for diverse response types.
The paper tackles matching a desired mean in causal systems through shift interventions.
problem Matching a desired mean in causal systems.
method Defining Markov equivalence classes, proposing active learning strategies, deriving lower bounds.
result Proposed active learning strategies require fewer interventions than previous approaches, especially for certain graph classes.