New method evaluates policies with latent confounders using optimal balance.
problem Evaluating policies with unobserved confounders in costly exploration scenarios.
method Importance weighting method to avoid latent outcome regression, minimizing adversarial balance objective.
result Provable consistency in policy evaluation with latent confounders, demonstrated empirically.
Enhances reinforcement learning with hierarchical policies using latent variables.
problem Improving performance in reinforcement learning tasks with hierarchical policies.
method Training each layer of a hierarchical neural network to solve tasks directly, with latent variables controlling lower layers.
result Improves performance on standard benchmark tasks and complex sparse-reward tasks.
New method infers latent policies from observations for imitation learning.
problem Learning from observation without expert actions.
method Characterizes causal effects and predicts likelihood of latent actions; uses action alignment for mapping.
result Corrected labeling of latent actions improves imitation performance.
Method adapts policies for new environments efficiently.
problem Difficulties in transferring reinforcement learning policies to new environments.
method Variational Policy Embedding (VPE) learns latent variables and a master policy.
result Policies can quickly adapt to new environments in latent space.
We decode latent states in Block MDPs and learn near-optimal policies.
problem Model estimation and reward-free learning in Block MDPs.
method Information-theoretical lower bound and efficient model estimation algorithm.
result Our algorithm approaches the information-theoretical limit for latent state decoding and converges to optimal policies.
Method trains multi-modal policy from unlabeled mixed demonstrations.
problem Training policies from unlabeled mixed demonstrations.
method Variational autoencoder with categorical latent variable to discover latent factors of variation.
result Policy can reproduce specific behaviors by conditioning on categorical vectors.
A novel approach for safe offline RL using latent safety constraints.
problem Balancing safety constraints and reward maximization in offline RL.
method Conditional Variational Autoencoders for latent safety modeling, Constrained Reward-Return Maximization.
result Our approach maintains safety compliance while optimizing rewards, outperforming existing methods.
We identify action representations from video data, proving their statistical benefits.
problem Identifying latent action policies from video data.
method Entropy-regularized LAPO objective, formalizing desiderata for action representations.
result Entropy-regularized LAPO identifies action representations satisfying desiderata under suitable conditions.
SeCTAR learns latent representations of trajectories for hierarchical reinforcement learning.
problem Learning lower layers in a hierarchy of reinforcement learning problems.
method SeCTAR uses variational autoencoders to learn latent representations of trajectories, combining policy and model consistency.
result SeCTAR effectively solves long-term and multi-stage problems with sparse rewards.
This research enhances exploration in DDPG using latent trajectory optimization.
problem Limited exploration in DDPG with deterministic policies.
method Model-based trajectory optimization for exploration in DDPG, using a learned deep dynamics model.
result Improved performance in continuous control tasks, especially with sparse rewards and images.
Adversarial attacks on probabilistic state-space models affect latent state and policy decisions.
problem Robust reinforcement learning under adversarial observability.
method Analyzing adversarial attacks on linear probabilistic state-space models.
result Demonstrating the influence of adversarial observations on latent state and policy decisions.
Deep reinforcement learning finds optimal learning policies for adaptive systems.
problem Finding individualized learning plans for learners with unknown latent traits.
method Formulated as a Markov decision process, applied deep Q-learning with a transition model estimator.
result The algorithm efficiently discovers optimal learning policies with small data sets.
Study tackles OPE in confounded settings, estimating policy value from proxies.
problem Difficulty in OPE due to unobserved confounders in infinite-horizon RL.
method Two-stage approach: estimating stationary distribution ratios and combining optimal balancing.
result Policy value can be identified from off-policy data with proxies and latent variable model.
A new method learns robust policies from offline data with latent structures.
problem Conservative policies under unrealistic dynamics shifts.
method d-RRMDP framework with f-divergence regularization and R2PVI algorithm. result R2PVI learns robust policies with superior computational efficiency.
New approach uses Wasserstein distances to score and optimize policy behaviors.
problem Comparing reinforcement learning policies and guiding policy optimization.
method Dual formulation of Wasserstein distances in latent behavioral space, learning score functions, smoothed WDs, stochastic gradient descent, on-policy algorithms.
result Demonstrated improved performance over existing methods in various environments.
New method uses variational inference to handle confounding in imitation learning.
problem Confounding due to different sensory inputs between expert and imitating agent.
method Train variational inference model to infer expert's latent information and use for latent-conditional policy training.
result Algorithm converges to correct interventional policy and achieves asymptotically optimal performance.
Method learns latent states from rich observations to improve RL exploration.
problem Improving RL performance with rich observations and latent states.
method Estimates latent states from observations through regression and clustering, providing finite-sample guarantees.
result Exponential improvement over Q-learning with naïve exploration. Improved exploration in RL with latent state marginalization.
problem Complexity of deep probabilistic models limits their practical use in reinforcement learning.
method Adopting latent variable policies within the MaxEnt framework, with low-cost marginalization of latent states.
result Effective marginalization leads to better exploration and more robust training.
This study models FOMC policy decisions using debate-based LLMs.
problem Accurately predicting central bank policy decisions, especially FOMC's, is challenging.
method A novel framework that simulates FOMC's collective decision-making process through iterative rounds of LLMs interacting as agents.
result The debate-based approach significantly outperforms standard LLMs in prediction accuracy.
CompILE learns reusable segments from demonstrations for hierarchical task execution.
problem Learning reusable, variable-length segments of hierarchical behavior from demonstrations.
method Unsupervised, fully-differentiable sequence segmentation module for latent encoding and re-composition.
result Model generalizes to longer sequences and unseen environments, learns task boundaries and event encodings.
GFlowNet-EM learns complex latent variable models with discrete structures.
problem Challenges in modeling posteriors over discrete compositional latents with expectation-maximization.
method Uses GFlowNets to learn stochastic policies for sampling from complex posterior distributions.
result GFlowNet-EM enables training expressive LVMs with discrete compositional latents.
In this work, we propose a method for learning driver models that account for variables that cannot be observed directly. When trained on a synthetic dataset, our models are able to learn encodings for vehicle trajectories that distinguish between four distinct classes of driver behavior. Such encodings are learned wit…
DSE learns transferable skills across changing dynamics and goals.
problem Learning transferable skills across different reinforcement learning tasks.
method Variational inference for multi-task reinforcement learning with shared and task-specific latent spaces.
result Policies can generalize to unseen dynamics and goals conditions.
New method recovers diverse policies from expert data using state-action pair weighting.
problem Recovering diverse policies from expert trajectories.
method Pointwise mutual information weighted behavioral cloning.
result Effective in focusing on state-action pairs most representative of the style.
Identifies latent actions and dynamics from offline data with diverse demonstrators.
problem Recovering latent actions and environment dynamics from action-free trajectories.
method Assumes distinct policies for each demonstrator, identifies latent transitions and policies via matrix factorization.
result Identifies latent transitions and demonstrator policies up to permutation.
New method learns near-optimal policies with polynomial samples in A and H.
problem Episodic latent MABs with partial observations are challenging.
method Experiment design and method-of-moments approach.
result Polynomial samples in A and H for near-optimal policy learning.
Paper proposes a method to optimize policies for diverse individuals using heterogeneous data.
problem Learning optimal policies for a heterogeneous population from pre-collected data.
method Individualized offline policy optimization framework for heterogeneous MDPs.
result The proposed P4L algorithm achieves a fast rate of average regret.
A new HRL method learns hierarchical policies using mutual information maximization.
problem Learning hierarchical policies in reinforcement learning for structured tasks.
method Mutual information maximization for latent variable learning, advantage-weighted importance sampling for option policies, deterministic policy gradient for optimization.
result Enhanced performance in continuous control tasks through learned hierarchical policies.
This work improves policy optimization by maximizing entropy of state distribution, leading to better exploration.
problem Lack of exploration in state space when maximizing policy entropy.
method Proposes maximizing the entropy of a lower bound approximation to the state weighting distribution, based on latent space representation.
result Entropy regularization based on marginal state distribution achieves superior state space coverage and better performance in various domains.
VLBM learns MDP transitions from limited data, improving OPE performance.
problem Limited coverage of state and action space in offline trajectories.
method VLBM uses variational inference with RSA and branching architecture.
result VLBM outperforms existing OPE methods on deep OPE benchmark.
Researchers study how teachers' advising relationships influence their perceptions of satisfaction and students, not policy influence.
problem Understanding the relationship between teachers' advising relationships and their perceptions of satisfaction and students.
method Proposed a novel joint model of network and item responses (JNIRM) with correlated latent variables.
result Teachers' advising relationships contribute more to satisfaction and students than to influence over educational policies.
Novel approach for sim-to-real transfer using MPC and task representations.
problem Difficulty of sim-to-real transfer systems producing generalizable policies.
method Model-predictive control (MPC) and task representation learning.
result Direct transfer of multi-skill policy to real robot for unseen tasks.
SLAC learns latent representations for image-based RL tasks.
problem Challenges in learning policies from high-dimensional image observations.
method SLAC separates representation learning and task learning, using a latent variable model.
result SLAC outperforms model-free and model-based methods in image-based control tasks.
New method for evaluating policies in complex decision-making models with hidden variables.
problem Evaluating policies in partially observable Markov decision processes with hidden confounders.
method Introduces novel identification methods and minimax estimation techniques for linking target policy's value and observed data distribution.
result Proposes three estimators for off-policy evaluation in POMDPs with latent confounders, demonstrating their effectiveness through nonasymptotic and asymptotic analysis.
Deep RL predicts car steering angles from images.
problem Learning steering angles for autonomous cars in simulators.
method Extracts latent representations, trains RL on latent vectors.
result Method learns steering angles without human control signals.
The paper learns robot skills from demonstrations without supervision.
problem Discovering robotic options from unlabelled demonstrations.
method Temporal variational inference for latent variable learning.
result The framework can learn options across multiple datasets.
New framework learns policies for partially observable systems.
problem Learning policies in partially observable dynamical systems.
method Partially Observable Bilinear Actor-Critic framework.
result Algorithm can learn against optimal policies in certain cases.
Paper introduces CageBO for optimizing complex public policy problems.
problem Complex decision-making and implicit constraints in public policy.
method CageBO framework using conditional variational autoencoder.
result CageBO outperforms baselines in optimizing large-scale police redistricting.
LC-SAC tackles non-stationary dynamics in reinforcement learning.
problem Degradation of deep RL methods in non-stationary environments.
method LC-SAC uses latent context encoders and contrastive loss for dynamic information capture.
result LC-SAC outperforms SAC on environments with drastic dynamics changes.
The paper proposes a method to transfer skills between tasks using disentangled latent policies.
problem Transfer learning in reinforcement learning struggles with diverse tasks without explicit supervision.
method Learning a small set of policies in a disentangled latent space that can be recombined to solve many tasks.
result Disentangled latent policies enable quick performance on many diverse tasks.
A new learning method for prosthetic arms without explicit rewards.
problem Learning a prosthetic arm to interact with users without explicit reward signals.
method Interaction-Grounded Learning, observing multidimensional context and feedback vectors, discovering latent reward signal.
result The algorithm can discover a latent reward signal and ground its policies for successful interaction.
MAVEN improves multi-agent exploration by hybridizing value and policy-based methods.
problem Exploration and suboptimality in complex multi-agent environments.
method MAVEN combines value and policy-based approaches with a latent space for hierarchical control.
result MAVEN achieves significant performance improvements on challenging multi-agent tasks.
Single episode policy transfer in RL without access to rewards.
problem Performing near-optimally in a single attempt with unknown dynamics.
method Optimizes a probe and inference model to estimate latent variables for universal control policy.
result Significantly outperforms existing adaptive approaches in diverse domains.
Develops RL algorithm for lifelong non-stationary environments.
problem Challenges of reinforcement learning in environments with persistent change.
method Formalizes lifelong non-stationarity, uses latent variable models, and leverages online learning and probabilistic inference.
result Substantial improvement in performance over non-reasoning approaches in lifelong non-stationary environments.
PBO framework optimizes latent preferences over multiple objectives.
problem Optimizing latent preferences with multiple conflicting objectives.
method Proposes DSTS, a multi-objective generalization of dueling Thompson sampling.
result DSTS outperforms benchmarks and provides asymptotic consistency.
Algorithm learns near-optimal policies for reward-mixing MDPs with few latent contexts.
problem Episodic reinforcement learning in reward-mixing Markov decision processes with a few latent contexts.
method Sample-efficient algorithm EM^2 using higher-order method-of-moments approach.
result Provides an ε-optimal policy using O(ε^(-2) * S^d A^d * poly(H, Z)^d) episodes for arbitrary M ≥ 2.
Representing a dialog policy as a recurrent neural network (RNN) is attractive because it handles partial observability, infers a latent representation of state, and can be optimized with supervised learning (SL) or reinforcement learning (RL). For RL, a policy gradient approach is natural, but is sample inefficient. I…
New method optimizes policies in non-stationary environments.
problem Optimizing policies in non-stationary, context-dependent environments.
method Two-phase approach: offline learning and online adaptation.
result Our method outperforms existing approaches in both synthetic and real-world datasets.