A new algorithm tackles delayed combinatorial semi-bandit with causal relations.
problem Optimizing decisions in a non-stationary environment with delayed and causally related rewards.
method Formalized as a non-stationary delayed combinatorial semi-bandit problem, the approach models causal relations with a directed graph in a stationary structural equation model. The agent learns these relations from delayed feedback to optimize decisions.
result Proved a regret bound for the proposed algorithm's performance.
Study adapts combinatorial semi-bandit for piecewise stationary, causally related rewards.
problem Nonstationary environment with changing base arms' distributions and causal relationships.
method Upper Confidence Bound (UCB) algorithm with change-point detector and group restart strategy.
result Regret upper bound reflecting effects of structural and distribution changes.
Adversarial CBO optimizes under interventions by adversaries and non-stationarities.
problem Optimizing in the presence of adversaries and non-stationary factors.
method Formalizes CBO as ACBO, introduces CBO-MW algorithm combining online learning and causal modeling.
result First algorithm with bounded regret for ACBO, achieving superior performance in synthetic and real-world environments.
Algorithm improves RL by discovering delayed causal relations.
problem Improving data-efficiency and interpretability in RL.
method Predicts observations with Markov assumption, introduces hidden variables to explain past events.
result Significantly improves RL performance on simulated and real tasks.
This work frames reward modelling from preferences as a causal problem.
problem Reward modelling from preference data for AI alignment.
method Causal inference approach to identify challenges and assumptions.
result Causally-inspired approaches improve model robustness.
Resolves spurious correlations in causal models via intervention design.
problem Spurious correlations lead to incorrect causal models in reinforcement learning environments.
method Proposes a method to design interventions that improve causal models by incentivizing agents to find errors.
result Experimental results show improved causal models compared to baselines.
New algorithm learns causal graph to minimize regret in bandits without full structure.
problem Learning optimal decisions in bandits with unknown causal graph and latent confounders.
method Two-stage approach: first learns ancestors and necessary confounders, second applies standard bandit algorithm.
result No full causal structure needed for optimal decisions; only necessary confounders are crucial.
This work explains RL policies using causal models, revealing important patterns and failures.
problem Understanding why RL policies succeed or fail in complex, high-dimensional systems.
method Developed a nonlinear Causal Model Reduction framework to learn simplified causal models from RL policy actions and rewards.
result The approach can uncover important behavioral patterns and failure modes in trained RL policies.
Causal KL improves on existing metrics for evaluating causal models.
problem Insufficient discrimination between causal models using edit-distance and KL divergence.
method Introducing Causal KL, an augmented KL divergence that considers causal relationships.
result Causal KL variants effectively distinguish between observationally equivalent models.
RL agents learn from a few tasks to generalize to new ones.
problem Creating efficient RL agents that can solve multiple tasks.
method GHP-MDPs model with latent variables for hidden parameters.
result State-of-the-art performance and sample-efficiency on new tasks.
GACBO optimizes unknown causal graphs with interventions.
problem Optimizing a target variable on an unknown causal graph with interventions.
method Graph Agnostic Causal Bayesian Optimisation (GACBO) seeks to balance exploitation and exploration of causal structures and functions.
result GACBO outperforms baselines in simulated and real-world applications.
A framework helps reinforcement learning agents understand and decompose tasks from human demonstrations.
problem Sparse-reward tasks where demonstrations are used as sources of causal knowledge.
method Develops causal models through observation and reasons from this knowledge to decompose tasks.
result A basic implementation of Reasoning from Demonstration (RfD) is effective in sparse-reward tasks.
A new IRL model recovers reward and state structure from expert demonstrations.
problem Limitation of classical maximum entropy model in capturing state structure.
method Generalized maximum causal entropy for IRL models.
result Empirically outperforms classical models in recovering reward and state structure.
The paper tackles causal bandits for SEMs, proposing algorithms that avoid estimating 2 N 2^N 2 N reward distributions.
problem Designing an optimal sequence of interventions in causal graphical models to minimize cumulative regret.
method Proposes two algorithms for causal bandits for linear structural equation models (SEMs), avoiding the estimation of 2 N 2^N 2 N reward distributions. result Cumulative regrets scale as i l d e O ( d L + 1 2 N T ) ilde{\cal O} (d^{L+\frac{1}{2}} \sqrt{NT}) i l d e O ( d L + 2 1 N T ) under bounded noise and parameter space. New algorithm finds causal parent nodes without graph learning.
problem Finding causal relationships without knowing the graph structure.
method Developed an efficient algorithm using atomic interventions.
result Algorithm optimally performs interventions with sublinear complexity.
CausalRM models rewards from user feedback, overcoming noise and bias.
problem Aligning language models with user preferences from noisy, biased feedback.
method Causal-theoretic reward modeling framework addressing noise and bias in observational feedback.
result CausalRM learns accurate reward signals from noisy and biased observational feedback.
The paper explores how to apply causal knowledge across different datasets to improve learning.
problem How to apply causal knowledge across different datasets to improve learning.
method Investigates the structural causal bandit with transportability, fusing priors from source environments to enhance learning in the deployment setting.
result Achieves a sub-linear regret bound with an explicit dependence on informativeness of prior data, potentially outperforming standard bandit approaches.
This paper identifies the minimal set of nodes for optimal conditional interventions in causal bandits.
problem Optimizing decision-making in causal bandits with conditional interventions.
method Graphical characterization and efficient algorithm to identify the minimal set of nodes.
result The proposed algorithm significantly prunes the search space and accelerates convergence rates.
The paper tackles bandit problems with biased offline data by using causal methods.
problem Improving bandit algorithms with biased offline data that includes confounding and selection biases.
method Formalizes the problem from a causal perspective, categorizes biases, and derives robust bounds for each arm.
result Causal bounds can guide the bandit agent to learn a nearly-optimal decision policy and consistently reduce asymptotic regret.
New formalism for decision making combines causal structures with MDPs, improving reinforcement learning performance.
problem Sequential decision making with causal knowledge to improve performance.
method Causal Markov Decision Processes (C-MDPs) and C-UCBVI algorithm exploiting causal structure.
result C-UCBVI achieves an i l d e O ( H S Z T ) ilde{O}(HS\sqrt{ZT}) i l d e O ( H S Z T ) regret bound, independent of actions. Fine-tuning LLMs with observational data can lead to spurious correlations, but DeconfoundLM can mitigate this.
problem Aligning LLMs with human preferences and business objectives using observational data.
method DeconfoundLM, a method that removes confounders from reward signals.
result DeconfoundLM improves recovery of causal relationships and mitigates spurious correlations.
Optimizes experiment design for causal structure learning in linear models with cycles.
problem Causal structure learning from combined observational and interventional data in linear non-Gaussian cyclic models.
method Combinatorial characterization of equivalence classes, adaptive stochastic optimization, greedy policy with near-optimal performance guarantee, sampling-based estimator for reward function.
result Optimal experiment design reduces the equivalence class of causal graphs to a single true graph with a small number of interventions.
Paper tackles transfer RL under unobserved context, developing methods to reduce bias.
problem Transfer RL with unobserved contextual information leading to biased models.
method Develops causal bounds on transition and reward functions using demonstrator's data.
result Proposes Q learning and UCB-Q learning algorithms that converge to true value function without bias.
In this work we define and study the relations between Lorentzian Manifolds given by the diffeomorphisms which map causal future directed vectors onto causal future directed vectors. This class of diffeomorphisms, called proper causal relations, contains as a subset the well-known group of conformal relations and are d…
Relational Structural Causal Models enable causal reasoning about unseen object combinations.
problem Developing a model that can reason about causal and combinatorial aspects of unseen object combinations.
method Relational Structural Causal Models extend structural causal models to include relational variables and define identification criteria.
result Proposed relational neural causal models outperform non-relational baselines on simulated traffic scenes.
Chronological Causal Bandits (CCB) tackles dynamic causal decision-making.
problem Dynamic causal decision-making in a system where rewards depend on past interventions.
method Introduces a new MAB problem (Chronological Causal Bandit) where rewards are influenced by a dynamic causal model.
result Early findings show the CCB can transfer information between sequential MABs.
New model tackles causal bandits with dependent variables.
problem Understanding reward-maximizing interventions in causal networks with dependent variables.
method Introduces hierarchical causal bandit model with a contextual variable capturing interactions among variables.
result Derives nearly matching regret bounds for binary context in causal bandits with dependent arms.
We propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents' actions. Causal influence is assessed using counterfactual reasoning. At each timestep, an agent simulates alternate actions t…
New algorithms for efficient causal interventions with budget constraints and without constraints.
problem Efficiently learning best interventions in causal graphs with budget constraints.
method Developed algorithms for both budgeted and non-budgeted causal bandits, optimizing regret and side-information usage.
result Proposed algorithms minimize cumulative regret and perform better than standard methods.
Study improves CI tests for relational data to robustly discover causal structures.
problem Learning causal relationships from relational data.
method Conduct CI tests against relational data to robustly recover causal structure.
result Effective approach demonstrated through experiments.
We define and study a new kind of relation between two diffeomorphic Lorentzian manifolds called {\em causal relation}, which is any diffeomorphism characterized by mapping every causal vector of the first manifold onto a causal vector of the second. We perform a thorough study of the mathematical properties of causal …
Algorithm learns causal structures from time-series data, reducing tests for temporal vs. contemporaneous relations.
problem Learning causal structures from time-series data with latent confounders.
method Constraint-based algorithm that refines a causal graph by learning temporal relations first, then contemporaneous ones.
result Reduces the number of statistical tests and improves accuracy for synthetic and real-world data.
Discovering and exploiting the causal structure in the environment is a crucial challenge for intelligent agents. Here we explore whether causal reasoning can emerge via meta-reinforcement learning. We train a recurrent network with model-free reinforcement learning to solve a range of problems that each contain causal…
Optimizes user marketing campaigns to balance cost and effectiveness.
problem Lack of methods to optimize marketing campaigns considering cost and effectiveness.
method Proposes a treatment effect optimization algorithm using deep learning to balance cost and effectiveness.
result Demonstrates superior performance in cost-efficiency and real-world business value.
The paper tackles causal bandits with unknown SCMs and soft interventions, providing upper and lower bounds on regret.
problem Optimizing interventions in a causal system with unknown SCMs and soft interventions.
method Assumes unknown SCMs from a general class, allows infinite interventions, and provides upper and lower bounds on regret.
result General upper and lower bounds on cumulative achievable regret for various SCMs.
New models infer causal effects from graph-based time-series data.
problem Inferring causal effects from graph-based relational time-series data.
method Proposes causal inference models leveraging graph topology and time-series data.
result Relational time-series causal inference models accurately estimate local causal effects of individual nodes.
Language helps RL agents learn complex relational and causal structures.
problem Learning relational and causal structure in complex environments.
method Training RL agents to predict language descriptions and explanations.
result Language aids agents in learning challenging relational and causal tasks.
Local method identifies causal relations in Markov equivalent DAGs.
problem Identifying causal relations when multiple DAGs are Markov equivalent.
method Graphical condition and local criteria for identifying causal paths.
result Local learning algorithm efficiently identifies causal variables.
We introduce an off-policy evaluation procedure for highlighting episodes where applying a reinforcement learned (RL) policy is likely to have produced a substantially different outcome than the observed policy. In particular, we introduce a class of structural causal models (SCMs) for generating counterfactual traject…
This work explores adaptive strategies for multi-armed bandits with causal structure, achieving optimal regret bounds.
problem Adapting to causal structure in multi-armed bandits with additional observed variables.
method Reduction to linear bandits and establishment of Pareto optimal frontier of adaptive rates.
result Established upper and lower bounds on adaptive rates, resolving open questions.
Amortized Causal Discovery learns to infer causal graphs from time-series data, improving performance.
problem Inference of causal graphs from time-series data is inefficient due to fitting new models for each sample.
method Proposes Amortized Causal Discovery, a variational model that leverages shared dynamics across samples with different causal graphs.
result Significant improvements in causal discovery performance demonstrated experimentally.
The study examines causal razors and their logical relations, highlighting a dilemma in causal discovery.
problem Selecting a reasonable scoring criterion for causal discovery algorithms.
method Review and logical comparison of numerous causal razors, focusing on parameter minimality in multinomial models.
result Parameter minimality poses a dilemma in selecting a reasonable scoring criterion for causal discovery algorithms.
Develops RL algorithm for non-Markovian, non-stationary reward streams.
problem Maximizing rewards from non-Markovian, non-stationary reward streams.
method Uses causal DAG to construct Markov states, solves periodic MDP.
result Optimal state construction maximizes discounted rewards.
New scoring rule predicts causal relations from data with selection bias.
problem Discovering causal relations from independence constraints under selection bias and confounding.
method Local Y-Structure patterns and a scoring rule for Y-Structures.
result Y-Structure scoring rule successfully predicts causal relations in real-world data.
New method identifies cause-effect relations in multivariate time series data.
problem Identifying cause-effect relations in multivariate time series data.
method Fictitious vector autoregressive model to identify long-run relations and causality strength.
result High accuracy in identifying true cause-effect relations in simulations and climate change analysis.
The paper tackles reward-relevance in offline RL with sparse decision dynamics.
problem Offline reinforcement learning with sparse decision dynamics and estimation sparsity.
method Reward-filtered least-squares policy evaluation using thresholded lasso.
result The method provides theoretical guarantees with sample complexity dependent on sparse component size.
New method for evaluating sequential recommendations with lower variance.
problem Evaluating good sequences of music, video, news, and e-commerce recommendations.
method Proposes a new counterfactual estimator for sequential reward interactions with lower variance and asymptotic unbiasedness.
result Our method outperforms existing methods in bias and data efficiency for sequential track recommendations.
Detect spacetime curvature with event causality measurements.
problem Detecting spacetime curvature without rulers and clocks.
method Prove spacetime non-flatness through causal relations.
result Sixteen measurements verify non-flatness of non-conformally flat spacetimes.