Identifies latent actions and dynamics from offline data with diverse demonstrators.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We identify action representations from video data, proving their statistical benefits.
In this paper, we describe a novel approach to imitation learning that infers latent policies directly from state observations. We introduce a method that characterizes the causal effects of latent actions on observations while simultaneously predicting their likelihood. We then outline an action alignment procedure th…
UWM-JEPA predicts future scenarios in belief space, improving accuracy in partially observed environments.
Model captures system input variations in latent space for actionable dynamics.
New algorithm tackles nonstationary linear bandits with latent dynamics.
Latent variable models improve RL by facilitating efficient learning and exploration.
A new method for disentangling action sequences improves model stability.
New method recovers diverse policies from expert data using state-action pair weighting.
Future autonomous systems need reliable world models and complex action sequences.
Developing a dialogue agent that is capable of making autonomous decisions and communicating by natural language is one of the long-term goals of machine learning research. Traditional approaches either rely on hand-crafting a small state-action set for applying reinforcement learning that is not scalable or constructi…
EbC learns equivariant embeddings from unlabeled group actions.
This research tackles intervention-centric causal reasoning in learning agents by using meta-learning.
Optimal control in latent factor models uses Tsallis entropy for exploration.
Efficient RL in large POMDPs with latent determinism and embeddings.
New model predicts drug effects across various cell types using causal imputation.
This work exploits action equivariance for representation learning in reinforcement learning. Equivariance under actions states that transitions in the input space are mirrored by equivalent transitions in latent space, while the map and transition functions should also commute. We introduce a contrastive loss function…
Study tackles OPE in confounded settings, estimating policy value from proxies.
We introduce Bayesian Poisson Tucker decomposition (BPTD) for modeling country--country interaction event data. These data consist of interaction events of the form "country took action toward country at time ." BPTD discovers overlapping country--community memberships, including the number of latent com…
New algorithms tackle latent bandit problems with lower regret.
Agent decides when to measure latent states in RL to improve efficiency.
Alpha signals for statistical arbitrage strategies are often driven by latent factors. This paper analyses how to optimally trade with latent factors that cause prices to jump and diffuse. Moreover, we account for the effect of the trader's actions on quoted prices and the prices they receive from trading. Under fairly…
In the present paper, we propose an extension of the Deep Planning Network (PlaNet), also referred to as PlaNet of the Bayesians (PlaNet-Bayes). There has been a growing demand in model predictive control (MPC) in partially observable environments in which complete information is unavailable because of, for example, la…
A new framework for structured bandits using influence diagrams and variational Thompson sampling.
New method identifies latent variables with sparse perturbations.
We introduce a method to design a computationally efficient -invariant neural network that approximates functions invariant to the action of a given permutation subgroup of the symmetric group on input data. The key element of the proposed network architecture is a new -invariant transformation modul…
We introduce a method for learning the dynamics of complex nonlinear systems based on deep generative models over temporal segments of states and actions. Unlike dynamics models that operate over individual discrete timesteps, we learn the distribution over future state trajectories conditioned on past state, past acti…
Latent-state environments with long horizons, such as those faced by recommender systems, pose significant challenges for reinforcement learning (RL). In this work, we identify and analyze several key hurdles for RL in such environments, including belief state error and small action advantage. We develop a general prin…
Paper learns meaningful state and action representations from MDP trajectories.
Modified Wasserstein metric for Gaussian distributions, invariant to isometries.
We present a dual-view mixture model to cluster users based on their features and latent behavioral functions. Every component of the mixture model represents a probability density over a feature view for observed user attributes and a behavior view for latent behavioral functions that are indirectly observed through u…
Intelligent agents can learn to represent the action spaces of other agents simply by observing them act. Such representations help agents quickly learn to predict the effects of their own actions on the environment and to plan complex action sequences. In this work, we address the problem of learning an agent's action…
We address the task of simultaneous feature fusion and modeling of discrete ordinal outputs. We propose a novel Gaussian process(GP) auto-encoder modeling approach. In particular, we introduce GP encoders to project multiple observed features onto a latent space, while GP decoders are responsible for reconstructing the…
New framework tackles stochastic latent subgroup heterogeneity in online decision-making.
New algorithm allows IGL to work with action-inclusive feedback.
In this work, we propose a method for learning driver models that account for variables that cannot be observed directly. When trained on a synthetic dataset, our models are able to learn encodings for vehicle trajectories that distinguish between four distinct classes of driver behavior. Such encodings are learned wit…
Deep Reinforcement Learning (RL) recently emerged as one of the most competitive approaches for learning in sequential decision making problems with fully observable environments, e.g., computer Go. However, very little work has been done in deep RL to handle partially observable environments. We propose a new architec…
VLBM learns MDP transitions from limited data, improving OPE performance.
Models for recommender systems use latent factors to explain the preferences and behaviors of users with respect to a set of items (e.g., movies, books, academic papers). Typically, the latent factors are assumed to be static and, given these factors, the observed preferences and behaviors of users are assumed to be ge…
Human players in professional team sports achieve high level coordination by dynamically choosing complementary skills and executing primitive actions to perform these skills. As a step toward creating intelligent agents with this capability for fully cooperative multi-agent settings, we propose a two-level hierarchica…
New method infers and samples point processes from latent diffusion.
Computer simulations have become a popular tool of assessing complex skills such as problem-solving skills. Log files of computer-based items record the entire human-computer interactive processes for each respondent. The response processes are very diverse, noisy, and of nonstandard formats. Few generic methods have b…
Paper models treatment effects by clustering patients with distinct survival characteristics.
Algorithm improves imitation learning from visual data in partially observable environments.
A variational autoencoder (VAE) derived from Tsallis statistics called q-VAE is proposed. In the proposed method, a standard VAE is employed to statistically extract latent space hidden in sampled data, and this latent space helps make robots controllable in feasible computational time and cost. To improve the usefulne…
The study uncovers latent capabilities of language models via causal representation learning.
New method identifies stable latent variables across different domains using weak distributional invariances.
We address the problem of learning hierarchical deep neural network policies for reinforcement learning. In contrast to methods that explicitly restrict or cripple lower layers of a hierarchy to force them to use higher-level modulating signals, each layer in our framework is trained to directly solve the task, but acq…