Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

58115173230 · Jun 202019922001200920182026
48 results for goal-directed exploration

GARIM theory explains how conscious manipulation of internal representations enhances goal-directed behavior.

problem Limited understanding of how consciousness supports flexible goal-directed cognition.
method Extending a three-component theory of flexible cognition, proposing GARIM theory.
result Conscious states actively manipulate internal representations to align with goals, enhancing flexibility.

UPNs embed planning within a goal-directed policy for effective visuomotor control.

problem Learning abstract representations for visuomotor control and generalization.
method Differentiable planning within a latent space, gradient descent trajectory optimization, end-to-end learning of representations.
result UPNs can transfer visuomotor planning strategies across robots with different morphologies and actuation capabilities.

Robot learns from multiple teachers to efficiently achieve various motor skill outcomes.

problem Efficiently learning motor skills from multiple teachers and strategies.
method Hierarchical active decisions based on empirical evaluation of learning progress.
result Significantly more efficient learning and coherent strategy selection.

This work improves Hindsight Learning for goal-directed tasks in reinforcement learning.

problem Sparse reward problems in reinforcement learning, especially goal-directed tasks.
method Improves Hindsight Experience Replay and Hindsight Policy Gradients for robotic tasks.
result Improved learning in Hindsight Experience Replay and Hindsight Policy Gradients.

Agent learns causal relationships from visual data to perform tasks.

problem Performing tasks in novel environments with latent causal structures.
method Learning-based approach to induce causal graphs from visual observations, using attention mechanisms.
result Effective generalization to new tasks with unseen causal structures.

Paper tackles goal-directed generation of discrete structures using conditional generative models.

problem Challenges in generating structured discrete data, especially for problems like program synthesis and materials design.
method Investigates conditional generative models to directly model the distribution of discrete structures given properties of interest. Introduces a novel approach to optimize a reinforcement learning objective.
result Improvements over maximum likelihood estimation and other baselines in generating molecules and identifying short python expressions.

Chance-constrained ActInf allows for small violations of constraints to drive goal-directed behavior.

problem Goal-directed behavior constrained by prior beliefs.
method Introducing chance constraints to ActInf, allowing for small violations of constraints.
result Chance-constrained ActInf allows for a trade-off between robust control and chance constraint violation.

PFAx extends PFA for global navigation in multiroom environments.

problem Global navigation in multiroom environments.
method SFA-based algorithm that decomposes tasks into subgoals, each solvable by PFAx.
result Stable global navigation in multiroom environments.

GCPN uses reinforcement learning to generate molecules optimizing desired properties.

problem Generating novel molecules with desired properties while obeying physical laws.
method Graph Convolutional Policy Network (GCPN) trained with reinforcement learning.
result GCPN achieves significant improvements in molecule optimization tasks.

Paper proposes Imitative Models combining IL and planning for flexible goal achievement.

problem Difficult to direct IL to arbitrary goals and specify reward functions for goal-directed planning.
method Imitative Models are probabilistic predictive models that plan interpretable expert-like trajectories.
result Imitative Models outperform IL and planning approaches in a dynamic autonomous driving task.

Universal AI seeks high-optionality states through empowerment and curiosity.

problem Understanding and optimizing AI behavior in uncertain environments.
method Unified framework combining AIXI and variational empowerment, showing how universal AI agents balance goal-directed behavior with uncertainty reduction curiosity.
result Self-AIXI asymptotically converges to AIXI performance and exhibits power-seeking behavior due to intrinsic motivations.

NIPA aims to translate brain learning mechanisms into scalable Bayesian inference.

problem Scalable Bayesian inference for large-scale statistical machine learning problems.
method Neural-inspired algorithm combining model-based, model-free, and episodic-control modules.
result Advances Bayesian methods and facilitates their application to deep learning.

Guided Learning improves end-to-end modeling for multi-stage decision-making.

problem Challenges in training unified neural networks for multi-stage decision-making.
method Guided Learning framework with a guide function and utility function.
result Significant improvement in performance over traditional methods.

New approach decouples skill learning and language grounding for autonomous agents.

problem Autonomous acquisition of skills without external instructions and feedback.
method Language-Goal-Behavior (LGB) architecture with semantic representation.
result Decouples skill learning and language grounding, enabling diversity and strategy switching.

DISTANA improves weather prediction by inferring hidden factors from temperature data.

problem Inferring hidden factors in spatiotemporal processes without supervision.
method Enhanced DISTANA architecture for spatiotemporal data, active tuning for latent state inference.
result DISTANA achieves more accurate predictions than other methods, inferring hidden factors from temperature data.

Active inference minimizes expected free energy for optimal behavior.

problem Understanding and optimizing behavior in complex systems.
method Combines Bayesian decision theory, optimal Bayesian design, and the free energy principle.
result Active inference emerges as a unified framework for information-seeking, utility maximization, and goal-directed behavior.

GraphAF generates chemically valid molecules efficiently and accurately.

problem Generating chemically valid molecular structures while optimizing chemical properties.
method Flow-based autoregressive model combining autoregressive and flow-based approaches.
result GraphAF generates 68% chemically valid molecules without chemical knowledge rules and 100% with rules, achieving state-of-the-art performance.

A neural network model mimics body functions for movement tasks.

problem Solving inverse and forward kinematics, dynamics for a redundant manipulator.
method Recurrent neural network with Mean of Multiple Computations principle, dynamic extension.
result Neural network solves inverse tasks and shows prototypical population-coding.

Unified framework for hybrid learning and optimization via active inference.

problem Sequential decisions in black-box evaluations requiring both task improvement and uncertainty reduction.
method Pragmatic Curiosity (PraC) framework that evaluates queries by balancing information gain and pragmatic value.
result Unified approach reduces decision risk and improves coverage of critical regions without task-specific rules.

Adapts Floyd-Warshall algorithm for RL to improve multi-goal task learning.

problem Limited transfer of learned information in model-free RL for dynamic goal tasks.
method Adapts Floyd-Warshall algorithm for RL to learn goal-conditioned action-value functions.
result FWRL achieves higher reward strategies in multi-goal tasks with fewer samples.

A new multi-objective RL framework improves intrinsic exploration performance.

problem Sub-optimal exploration performance due to ad-hoc handling of intrinsic exploration.
method A multi-objective RL framework where both exploration and exploitation are optimized as separate objectives.
result EMU-Q method outperforms classic and other intrinsic RL methods on benchmarks.

Study explores how to efficiently explore communities with limited budget.

problem Maximizing the number of members met with limited budget in community exploration.
method Systematic study from offline optimization to online learning, including greedy methods and upper confidence algorithms.
result Achieved logarithmic and constant regret bounds in online learning setting.

This work learns latent representations to speed up exploration in complex environments.

problem Challenging exploration in high-dimensional state and action spaces with sparse rewards.
method Representation learning using prior experience to learn effective latent representations.
result Learned latent representations reduce the dimensionality of the search space for effective exploration.

Go-Explore improves performance on hard-exploration problems in Atari games.

problem Challenges in reinforcement learning, especially with sparse or deceptive rewards.
method Exploits principles of remembering states, returning to promising states, and solving simulated environments.
result Scores significantly higher than previous state-of-the-art on Montezuma's Revenge and Pitfall.

This work tackles the exploration-exploitation dilemma in RL by developing optimal policies that are inherently exploration-conscious.

problem The exploration-exploitation tradeoff in Reinforcement Learning, where policies need to balance new action exploration with past experience exploitation.
method Developed exploration-conscious criteria that result in optimal policies, solving these criteria by solving a surrogate Markov Decision Process.
result Demonstrated superior performance of exploration-conscious RL algorithms compared to non-exploration-conscious counterparts in both discrete and continuous action spaces.

R3L uses planning algorithms to efficiently explore sparse reward environments.

problem Balancing exploration and exploitation in sparse reward reinforcement learning.
method Formulate exploration as a search problem using RRT, leverage demonstrations from initial solutions to refine RL policy.
result R3L outperforms classic and intrinsic exploration techniques, requiring fewer samples and achieving better asymptotic performance.

This work proposes sample complexity bounds for Q-learning with random exploration.

problem Understanding the sample efficiency of simple exploration strategies in reinforcement learning.
method Problem-specific sample complexity bounds for Q-learning with random walk exploration.
result Proposes bounds that relate to empirical performance in benchmark domains.

New exploration bonuses improve reinforcement learning efficiency.

problem Efficient exploration in unknown environments with limited feedback.
method Improved exploration bonuses scaling with 1/n and improved stopping time analysis.
result Faster learning rates and improved sample complexity in pure-exploration settings.

New findings on when to use action space exploration in reinforcement learning.

problem Understanding when to use action space exploration over traditional methods.
method Theoretical analysis and empirical testing of simple exploration methods.
result Exploration in action space is preferred when parametric complexity exceeds action space dimensionality and horizon length.