This work improves reinforcement learning with sparse rewards by following diverse past trajectories.
problem Challenges in reinforcement learning with sparse rewards and myopic behavior.
method Proposes a trajectory-conditioned policy to learn from a memory buffer of diverse past trajectories.
result Significantly outperforms existing methods on complex tasks with local optima.
New method improves sample diversity and efficiency from complex distributions.
problem Sampling from intractable un-normalized distributions with high auto-correlation.
method Stein self-repulsive dynamics using a repulsive force to push samples away from past trajectories.
result Significantly decreases auto-correlation and increases effective sample size.
New method recovers diverse policies from expert data using state-action pair weighting.
problem Recovering diverse policies from expert trajectories.
method Pointwise mutual information weighted behavioral cloning.
result Effective in focusing on state-action pairs most representative of the style.
The paper clusters hypergraphs to find diverse and experienced groups based on past experiences.
problem Finding diverse and experienced groups with respect to past experiences.
method Regularized edge-based hypergraph clustering objective with a 2-approximation algorithm.
result Demonstrates an efficient 2-approximation algorithm for clustering hypergraphs.
We introduce a method for learning the dynamics of complex nonlinear systems based on deep generative models over temporal segments of states and actions. Unlike dynamics models that operate over individual discrete timesteps, we learn the distribution over future state trajectories conditioned on past state, past acti…
CoverNet predicts urban driving trajectories using diverse sets of possible actions.
problem Multimodal probabilistic trajectory prediction for urban driving.
method Frame trajectory prediction as classification over a diverse set of trajectories; dynamically generate sets based on current state.
result Outperforms state-of-the-art methods on real-world self-driving datasets.
PhysVarMix predicts diverse urban trajectories with physics constraints.
problem Predicting complex urban agent trajectories with multiple plausible scenarios.
method Physics-informed variational mixture model combining learning and physics constraints.
result Superior performance compared to existing methods on benchmark datasets.
This paper explores a simple regularizer for reinforcement learning by proposing Generative Adversarial Self-Imitation Learning (GASIL), which encourages the agent to imitate past good trajectories via generative adversarial imitation learning framework. Instead of directly maximizing rewards, GASIL focuses on reproduc…
This work uses MMD to find diverse policies in reinforcement learning.
problem Finding multiple near-optimal policies for a task.
method Formalizes policy difference as trajectory distribution discrepancy, uses MMD for optimization.
result Derives gradient-based optimization for diverse policy identification.
Improved Gaussian Process model for predicting trajectories without independence assumption errors.
problem Incorrect independence assumption in previous work on Gaussian Process uncertainty propagation.
method Proposed a novel piecewise linear approximation to correct the independence assumption in continuous models.
result Corrected the independence assumption in Gaussian Process models for predicting trajectories.
End-to-end model predicts multiagent trajectories using game theory and neural nets.
problem Predicting trajectories of interacting agents in complex scenarios.
method Hybrid neural net with game-theoretic reasoning, using implicit layers to map preferences to Nash equilibria.
result Trains an interpretable model that predicts future trajectories and transfers to decision making.
Algorithm improves sample efficiency for dynamic task adaptation.
problem Sample inefficiency in imitation learning for changing task distributions.
method Assigns importance weights to past demonstrations for meta adaptation.
result Robot adapts to unseen environments with few demonstrations.
Improves diversity of text-to-image models without sacrificing FID.
problem Lack of diversity and tendency to recreate training set images.
method Adds sparse repellency terms to diffusion SDE to guide trajectories away from a reference set.
result Improves diversity of diffusion models with minimal impact on FID.
Representation learning of pedestrian trajectories transforms variable-length timestamp-coordinate tuples of a trajectory into a fixed-length vector representation that summarizes spatiotemporal characteristics. It is a crucial technique to connect feature-based data mining with trajectory data. Trajectory representati…
Safe active learning for time-series models with Gaussian processes.
problem Learning time-series models while respecting safety constraints.
method Employing Gaussian processes with a nonlinear exogenous input structure, the approach dynamically explores the input space to generate data for model learning.
result The approach effectively learns time-series models under safety constraints, as demonstrated in a technical application.
WayDCM predicts trajectories considering long-term goals, improving accuracy.
problem Predicting future trajectories of dynamic agents in complex environments.
method WayDCM combines DCM and NN to predict intermediate goals and trajectories, considering long-term goals.
result WayDCM outperforms previous methods on the Waymo Open dataset.
Training task diversity improves ICL with linear attention.
problem Understanding the impact of training task diversity on in-context learning.
method Modeling training task vectors as a mixture of low-rank Gaussians.
result Our model explains why training with task diversity shortens the ICL plateau and achieves out-of-distribution generalization.
Paper tackles continual reinforcement learning challenges with diversity exploration and adversarial self-correction.
problem Challenges in learning new tasks sequentially due to catastrophic forgetting.
method Develops an end-to-end framework (CDAN) combining unsupervised diversity exploration and adversarial self-correction.
result Final result outperforms baseline by 18.35% in NSD and 0.61 in average reward.
Analyzing the temporal behavior of nodes in time-varying graphs is useful for many applications such as targeted advertising, community evolution and outlier detection. In this paper, we present a novel approach, STWalk, for learning trajectory representations of nodes in temporal graphs. The proposed framework makes u…
Proposes an RL algorithm for handling convex constraints.
problem Handling constraints in reinforcement learning.
method Algorithm for convex constraints in RL.
result Algorithm can enforce new properties like diversity.
Improved GFlowNets learn more efficiently with trajectory balance.
problem Inefficient credit assignment in GFlowNets leads to suboptimal learning.
method Proposed trajectory balance as a new learning objective.
result Trajectory balance leads to more efficient and robust GFlowNet learning.
Predicts multiple vehicle trajectories efficiently.
problem Predicting uncertain future motions of agents in dynamic scenes.
method Probabilistic framework learning latent variables for multi-step future modeling.
result State-of-the-art predictions on vehicle trajectory datasets.
Identifies latent actions and dynamics from offline data with diverse demonstrators.
problem Recovering latent actions and environment dynamics from action-free trajectories.
method Assumes distinct policies for each demonstrator, identifies latent transitions and policies via matrix factorization.
result Identifies latent transitions and demonstrator policies up to permutation.
New framework predicts diverse, contextually plausible 3D human motions.
problem Predicting multiple plausible future 3D poses given observed poses.
method Developed a new variational framework that conditions latent variable on past observation to encourage relevant information.
result Our approach generates motions of higher quality and preserves contextual information.
New method for predicting paths of unpredictable objects with high confidence.
problem Need for dependable uncertainty estimates in motion planning with diverse unpredictable objects.
method Blend online conformal prediction, multiple time series techniques, and heteroscedasticity addressing.
result Simultaneous forecasting bands that cover entire paths with high probability.
Enhanced decision-making through Dreamer's anticipatory trajectories and Online Decision Transformer.
problem Efficiently integrating world models with decision transformers.
method Combining Dreamer's trajectory forecasting with Online Decision Transformer's adaptive learning.
result Notable improvements in sample efficiency and reward maximization.
The tremendous growth of positioning technologies and GPS enabled devices has produced huge volumes of tracking data during the recent years. This source of information constitutes a rich input for data analytics processes, either offline (e.g. cluster analysis, hot motion discovery) or online (e.g. short-term forecast…
Measuring similarities between unlabeled time series trajectories is an important problem in domains as diverse as medicine, astronomy, finance, and computer vision. It is often unclear what is the appropriate metric to use because of the complex nature of noise in the trajectories (e.g. different sampling rates or out…
RNN operators solve Newton's equations with large timesteps for molecular dynamics.
problem Solving Newton's equations of motion with large timesteps for molecular dynamics simulations.
method Recurrent Neural Networks (RNN) operators to solve Newton's equations using past trajectory data.
result Significant speedup in molecular dynamics simulations with timesteps up to 4000 times larger.
Dynamic treatment effects estimated over time using covariate balancing.
problem Estimating treatment effects in panel data with dynamic treatments.
method Dynamic covariate balancing with potential local projections.
result Established inferential guarantees for the proposed method.
Meta-Q-Learning improves meta-Reinforcement Learning using past data.
problem Improving meta-Reinforcement Learning performance.
method Meta-Q-Learning combines Q-learning, multi-task objectives, and off-policy updates to adapt policies from past data.
result Meta-Q-Learning compares favorably with state-of-the-art meta-RL algorithms on benchmarks.
CGNS predicts probabilistic trajectories for safer autonomous systems.
problem Accurate probabilistic trajectory prediction for dynamic obstacles in complex scenarios.
method CGNS combines latent space learning and variational divergence minimization, incorporating static and interaction information with soft attention mechanisms and regularization for soft constraints.
result CGNS outperforms baseline approaches in pedestrian trajectory prediction and naturalistic driving datasets.
Coordination recognition and subtle pattern prediction of future trajectories play a significant role when modeling interactive behaviors of multiple agents. Due to the essential property of uncertainty in the future evolution, deterministic predictors are not sufficiently safe and robust. In order to tackle the task o…
New method improves meta-reinforcement learning efficiency.
problem Sample inefficiency in meta-reinforcement learning.
method Hindsight Foresight Relabeling (HFR) method.
result HFR improves performance on various meta-reinforcement learning tasks.
Paper examines adversarial attacks on weather forecasting models, focusing on TC trajectory prediction.
problem Adversarial attacks can mislead downstream TC trajectory predictions in DLWF models.
method Proposes Cyc-Attack, a method using a surrogate model and skewness-aware loss function to generate adversarial TC paths.
result Cyc-Attack achieves higher true positive rates and lower false alarm rates compared to conventional methods.
MxPool learns graph features from diverse graphs using a hierarchical structure.
problem Learning graph features from diverse graphs with varying properties and sizes.
method MxPool uses a multiplex structure with multiple graph convolution/pooling networks in a hierarchical learning structure.
result MxPool outperforms state-of-the-art methods on graph classification benchmarks.
DISCO predicts system states from short trajectories using an evolved operator.
problem Predicting next states of dynamical systems governed by unknown PDEs.
method DISCO uses a hypernetwork to generate parameters of a smaller operator network for state prediction.
result DISCO achieves state-of-the-art performance with fewer training epochs and generalizes well.
New method learns influential action sequences without privileged final states.
problem Learning meaningful action sequences in large action spaces.
method Model-free approach that considers the full trajectory.
result Successfully applied to large action spaces.
Deep learning predicts tropical cyclone tracks efficiently.
problem Forecasting tropical cyclone trajectories with high precision and speed.
method Fused neural network model using past trajectory data and reanalysis atmospheric images.
result Deep learning can provide valuable and complementary predictions for tropical cyclone tracks.
A simple but useful method of reciprocal values is introduced, explained and illustrated. This method simplifies the analysis of hyperbolic distributions, which are causing serious problems in the demographic and economic research. It allows for a unique identification of hyperbolic distributions and for unravelling co…
EDU method finds diverse optimal solutions for expensive simulators.
problem Optimizing expensive black-box simulators for diverse solutions.
method EDU method searches for diverse locally-optimal solutions within a tolerance level.
result EDU yields a closed-form acquisition function facilitating efficient sequential queries.
New method improves learning from multiple correlated data trajectories.
problem Learning from multiple correlated data trajectories without mixing assumptions.
method Hellinger localization framework for maximum likelihood estimation.
result Instance-optimal bounds that scale with full data budget under broad conditions.
Model predicts multi-agent trajectories using a differentiable simulator.
problem Predicting future positions of multiple interacting agents.
method Conditional recurrent variational neural networks (CVRNNs) with a kinematic bicycle model.
result Achieves state-of-the-art results on INTERACTION dataset.
Recently, researchers proposed various low-precision gradient compression, for efficient communication in large-scale distributed optimization. Based on these work, we try to reduce the communication complexity from a new direction. We pursue an ideal bijective mapping between two spaces of gradient distribution, so th…
Presents STRIPE model for probabilistic forecasting of non-stationary time series.
problem Probabilistic forecasting of non-stationary time series.
method STRIPE model representing structured diversity based on shape and time features, with diversification mechanism using determinantal point processes (DPP).
result STRIPE significantly outperforms baseline methods for representing diversity while maintaining forecasting accuracy.
Novel framework predicts brain biomarker trajectories with superior performance.
problem Challenges in estimating longitudinal brain biomarker trajectories due to variability, inconsistencies, and irregular measurements.
method Personalized deep kernel regression with Adaptive Shrinkage Estimation.
result Superior predictive performance compared to state-of-the-art models.
The MoN loss fails to accurately represent ground truth probability density functions in probabilistic trajectory prediction.
problem Improving the diversity of probabilistic trajectory predictions in autonomous driving and robot planning.
method Proof and validation of the MoN loss's inaccuracy and proposed solutions to correct it.
result The MoN loss approximates the square root of the ground truth probability density function, not the function itself.
Algorithm learns actions from past states in complex tasks.
problem Learning policies from human feedback is expensive.
method Combining learned feature encoder with inverse models to simulate past actions.
result Algorithm can infer specific skills from single state.