Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

146292438584 · Jun 202019922001200920182026
48 results for space environment

Curiosity-driven exploration learns disentangled goal spaces for efficient complex environments.

problem Efficient exploration in complex environments with high-dimensional continuous actions.
method Learned disentangled goal spaces to reflect the environment's structure and maximize learning progress.
result Disentangled goal spaces lead to better exploration performances than entangled goal spaces.

OBSER framework infers sub-environments from objects, outperforming scene-based methods.

problem Zero-shot recognition of environments from object distributions.
method Bayesian framework using metric and self-supervised learning models to estimate object distributions in latent space.
result OBSER framework reliably performs inference in open-world and photorealistic environments, outperforming scene-based methods.

This paper introduces CENIE to quantify environment novelty for better UED.

problem Challenges in measuring environment novelty for effective UED.
method CENIE framework using state-action space coverage and Gaussian Mixture Models.
result CENIE improves UED performance across multiple benchmarks.

Adversarial RL recovers agent rewards from financial market data simulations.

problem Recovering agent rewards in volatile financial markets with unknown dynamics.
method Adversarial inverse reinforcement learning in latent space simulations.
result Adversarial RL can robustly recover agent rewards from latent space representations of real market data.

Generative model learns spatial memory for predicting agent's future in complex environments.

problem Training generative and temporal models in partially observed, 3D environments is challenging.
method Action-conditioned generative model with non-parametric spatial memory and state-space model for dynamics.
result Scalable architecture capable of coherent predictions over hundreds of time steps in 2D and 3D environments.

Research shows collective learning across diverse environments is hard due to privacy and security concerns.

problem Privacy, security, and equity concerns restrict information sharing in diverse AI environments.
method Characterized learning algorithms as choice correspondences, provided minimum requirements for rational learning algorithms.
result The only rational learning algorithm in heterogeneous environments is unilaterally learning from a single environment without information sharing.

SMEs provide a transparent testbed for RL evaluation.

problem Lack of precise, white-box diagnostics in RL environments.
method Synthetic Monitoring Environments (SMEs) with fully configurable task characteristics and known optimal policies.
result SMEs allow for precise evaluation of RL algorithms, revealing the impact of specific environmental properties.

A framework for learning disentangled representations of symmetric environments.

problem Discovering and modelling the underlying structure of environments.
method Group representation theory for disentangled representations of dynamical environments.
result Our method enables accurate long-horizon predictions and correlates with disentanglement quality.

A new method optimizes in nonstationary environments with many arms efficiently.

problem Optimizing in nonstationary environments with a large number of arms.
method Gaussian interpolation to learn continuous Lipschitz reward functions in nonstationary environments.
result Efficiently learns continuous Lipschitz reward functions with O(T)\mathcal{O}^*(\sqrt{T}) cumulative regret.

Proposes a method for model-based RL in complex environments without perfect simulators.

problem Lack of cheap and perfect simulators in real-world tasks.
method Induces a world program by learning dynamics and actions in graph-based environments.
result World program enables complex planning tasks in environments without perfect simulators.

Teacher algorithm helps DRL learn diverse environments efficiently.

problem Teach DRL to learn in various, unknown environments efficiently.
method Transformed into a bandit problem, learns to sample environments.
result ALP-GMM models learning progress, improving curriculum design.

A new exploration method for RL using parameter space noise.

problem Improving exploration in deep reinforcement learning.
method Switching isotropic and directional exploration in parameter space with parameter space noise.
result The proposed method achieves competitive results and better performance in sparse reward environments.

Deep belief networks improve Dyna-style planning in large state spaces.

problem Lack of real data and difficulty in learning a good generative model for large state spaces.
method Used deep belief networks to learn an environment model for Dyna-style planning.
result Deep belief networks significantly outperform linear expectation models in empirical validation.

Paper introduces CRLMaze, a new benchmark for continual reinforcement learning in 3D non-stationary environments.

problem Challenges of training reinforcement learning agents in high-dimensional, always-changing environments.
method End-to-end model-free continual reinforcement learning strategy.
result Competitive results in a complex 3D non-stationary task, outperforming four baselines.

DeepMDP simplifies complex observations into continuous latent states.

problem Learning from high-dimensional observations in reinforcement learning.
method Trains a DeepMDP model that predicts rewards and next latent states.
result Optimization of DeepMDP objectives ensures quality of latent space and environment model.

The paper tackles causal discovery and forecasting in nonstationary environments using state-space models.

problem Challenges in identifying causal relations and forecasting in nonstationary time series.
method Exploiting a particular type of state-space model to represent nonstationary processes, allowing changes in causal strengths and noise variances.
result Nonstationarity helps identify causal structure and improves forecasting.

RL agents optimize only specified features; this project infers unmentioned preferences from the state of the environment.

problem RL agents are indifferent to features not specified in a reward function, leading to unconsidered preferences.
method Developed an algorithm based on Maximum Causal Entropy IRL to infer preferences and side effects from the state of the environment.
result Information from the initial state can infer both side effects to avoid and preferences for environment organization.

This work uses reinforcement learning to optimize task scheduling and execution in a dynamic multi-agent warehouse environment.

problem Optimizing task scheduling and execution in a dynamic multi-agent warehouse environment with limited observability.
method Deep reinforcement learning to solve both high-level scheduling and low-level multi-agent execution problems.
result Demonstrates the effectiveness of reinforcement learning in optimizing task scheduling and execution in a dynamic multi-agent environment.

CEA augments reinforcement learning by generating counterfactual experiences.

problem Challenges in reinforcement learning, especially out-of-distribution and inefficient exploration.
method CEA uses variational autoencoders to model state transitions and introduces randomness for non-stationarity. It expands learning data through counterfactual inference.
result CEA outperforms SOTA algorithms in diverse environments.

HashReward improves imitation learning in high-dimensional environments by balancing reward generation and dimensionality reduction.

problem Making policies generalize well in high-dimensional state-action spaces, especially in game playing with raw pixel inputs.
method HashReward uses supervised hashing to balance reward generation and dimensionality reduction.
result HashReward outperforms state-of-the-art methods in high-dimensional environments.

KRCD detects unobserved confounders in nonlinear observational data.

problem Detecting unobserved confounders in nonlinear observational studies.
method Kernel Regression Confounder Detection (KRCD) using reproducing kernel Hilbert spaces.
result KRCD outperforms existing methods and achieves superior computational efficiency.

Develops HMRL for sparse reward RL problems, improving meta policy efficiency and transferability.

problem Difficulty in learning meta policies for sparse reward RL problems.
method Hyper-Meta RL framework with cross-environment meta state embedding and shaped meta reward.
result Improves meta policy generalization and efficiency for sparse reward RL problems.

Robots learn spatial perception from sensorimotor invariants.

problem Developing autonomous robots that perceive space without human intuition.
method Study how a robot's motor commands relate to changes in exteroceptive inputs to deduce its spatial configuration.
result Robots can learn the configuration space of their sensors, revealing a planar position and orientation.

Quantum algorithms speed up reinforcement learning policies in large state-action spaces.

problem Limitations of quantum access in training reinforcement learning policies.
method Designing quantum algorithms to train reinforcement learning policies.
result Quantum algorithms offer full quadratic speed-ups in sample complexity for well-behaved policies.

HOMER learns latent states to explore rich environments efficiently.

problem Exploration in rich observation environments with unknown latent states.
method Interleaves representation learning and strategic exploration to identify kinematic states.
result Provably efficient exploration with polynomial sample complexity in latent states and time horizon.