A new framework handles hybrid action spaces in reinforcement learning.
problem Handling reinforcement learning with both discrete and continuous actions.
method Parametrized Deep Q-Networks (P-DQN) framework integrating DQN and DDPG.
result Empirical validation of efficiency and effectiveness in simulated RoboCup soccer and game King of Glory.
Alpha Zero adapts to continuous action spaces for real-world tasks.
problem Real-world reinforcement learning domains often have continuous action spaces.
method Interleaves tree search and deep learning, extending Alpha Zero for continuous action spaces.
result Preliminary experiments on the Pendulum task show feasibility of the approach.
AQL uses amortized inference to handle high-dimensional action spaces in Q-learning.
problem Difficulty in maximizing over large action spaces in Q-learning.
method Replace expensive maximization over all actions with a maximization over a small subset sampled from a learned proposal distribution.
result AQL outperforms existing methods on continuous control tasks with up to 21 dimensional actions.
Smooth approximations for continuous functions on orbit spaces.
problem Approximating continuous functions on orbit spaces.
method Study of subcartesian spaces and proper Lie group actions.
result Continuous functions can be approximated by smooth functions.
FSQ algorithm extends Q-learning to continuous actions with linear complexity.
problem Extending Q-learning to continuous action spaces with linear complexity.
method Discretization of the action space to maintain linear complexity.
result FSQ algorithm achieves linear complexity in the discretized problem.
Multi-task learning improves robotic control in continuous action spaces.
problem Robotic control in continuous action spaces lacks effective multi-task learning methods.
method Applied multi-task learning methods to continuous action spaces and compared performance with baselines.
result Multi-task learning outperforms baselines and alternative methods in continuous control tasks.
Proper actions on bornological spaces are characterized with compatible coarse structures.
problem Characterizing proper actions on bornological spaces.
method Proving the existence of compatible coarse structures for proper actions.
result Bornological spaces admit compatible coarse structures for proper actions.
Novel framework proves fast RL convergence in continuous spaces.
problem Analyzing stability in continuous state-action RL.
method Introduces a novel framework to analyze stability properties of RL.
result Highlights two key stability properties and demonstrates their satisfaction in RL.
Efficient algorithms for contextual bandits with smooth regret in continuous action spaces.
problem Efficient learning in large or continuous action spaces.
method Smooth regret notion and efficient algorithms for general function approximation.
result Statistically and computationally efficient algorithms for contextual bandits with smooth regret.
The paper extends Thompson Sampling to infinite action spaces using information theory.
problem Addressing the limitation of finite action spaces in Thompson Sampling.
method Information-theoretic analysis, extending rate-distortion theory to infinite action spaces.
result Derives a near-optimal regret bound for bandits with infinite and continuous action spaces.
New proof shows path-connectedness of actions on intervals and circles.
problem Path-connectedness of C1+ac actions of Zd. method New proof using C1 diffeomorphisms with absolutely continuous derivative. result Path-connectedness of the space of actions.
A new policy gradient estimator reduces variance for clipped actions in continuous control tasks.
problem Policy gradient methods struggle with bounded action spaces.
method Proposes a new policy gradient estimator that accounts for clipped actions.
result The new estimator achieves lower variance and outperforms conventional methods.
New method extends low-rank MDPs to continuous action spaces.
problem Limited applicability of current low-rank MDP methods to continuous action spaces.
method Extending FLAMBE algorithm to continuous action spaces with Hölder smoothness conditions.
result Similar PAC bound achieved for continuous actions with polynomial dependence on smoothness order.
The paper tackles contextual bandits with continuous actions using smoothing and zooming techniques.
problem Learning with continuous action spaces in the context of contextual bandits.
method The approach involves smoothing and zooming techniques to handle the continuous action space and unknown smoothness parameters.
result Improved regret bounds and adaptive algorithms for contextual bandits with continuous actions.
Hybrid Policy Optimization tackles reinforcement learning in hybrid spaces, improving performance over PPO.
problem Credit assignment issues and biased gradients in hybrid discrete-continuous action spaces.
method Mixed gradient estimator combining pathwise and score-function gradients, reformulating problems in hybrid form.
result HPO substantially outperforms PPO on inventory control and switched systems, with performance gaps increasing with continuous action dimension.
Proposes hybrid reinforcement learning for both discrete and continuous control problems.
problem Real-world control problems involving both discrete and continuous decision variables.
method Solves hybrid problems by optimizing for discrete and continuous actions simultaneously.
result Efficiently solves hybrid reinforcement learning problems and improves upon expert heuristics.
Hybrid RL method optimizes trading by balancing continuous and discrete actions.
problem Optimal execution in algorithmic trading with continuous-discrete action space.
method Combines continuous and discrete RL agents for better trading decisions.
result Significantly outperforms existing methods in trading efficiency and stability.
New algorithms tackle multi-agent problems with hybrid action spaces.
problem Applying deep reinforcement learning to multi-agent problems with discrete-continuous hybrid action spaces.
method Proposed two novel algorithms: Deep MAPQN and Deep MAHHQN, using centralized training and decentralized execution.
result Empirical results show both algorithms significantly outperform existing methods.
Bootstrap policies improve regret in continuous state-action reinforcement learning.
problem Improving regret in reinforcement learning for continuous state and action spaces.
method Bootstrap-based policies for stochastic linear systems with quadratic cost functions.
result Bootstrap policies achieve a square root scaling of regret with respect to time.
Practical algorithm for contextual bandits with large action spaces.
problem Efficient algorithms for decision making in large, continuous action spaces.
method Uses computational oracles for supervised learning and optimization over the action space.
result Achieves sample complexity, runtime, and memory independent of the size of the action space.
Paper tackles RL with continuous actions and unmeasured confounders.
problem Offline policy learning with continuous actions and unmeasured confounders.
method Developed a novel identification result and a minimax estimator for nonparametric policy value estimation.
result Introduced a policy-gradient-based algorithm to identify the optimal policy.
Policy Prediction Network improves continuous control problems with model-free and model-based learning.
problem Improving sample complexity and performance in continuous control problems.
method Integrates model-free and model-based reinforcement learning, introduces implicit model-based learning for continuous action space.
result First to introduce implicit model-based learning to Policy Gradient algorithms for continuous action space.
It has long been assumed that high dimensional continuous control problems cannot be solved effectively by discretizing individual dimensions of the action space due to the exponentially large number of bins over which policies would have to be learned. In this paper, we draw inspiration from the recent success of sequ…
Bayes-CPACE optimally explores continuous BAMDPs.
problem Model uncertainty in continuous state and action spaces.
method Covering state-belief-action space with samples, exploiting Lipschitz continuity.
result Near-optimal value function computed efficiently.
This paper improves RL for PDE control with function-valued actions.
problem Control of PDEs with high-dimensional, spatially related actions.
method Action descriptors and deep deterministic policy gradient.
result Action descriptor approach is more sample efficient.
New algorithms learn MDPs with continuous states and actions using Gaussian processes.
problem Online learning in unknown, episodic MDPs with continuous states and actions.
method Developed variants of UCRL and posterior sampling algorithms using Gaussian process priors.
result Sublinear regret bounds for learning MDPs with specific kernel structures.
Proposes a value-based method for continuous control without an actor.
problem Computational infeasibility of evaluating Q-values in continuous action spaces.
method Structurally maximizable Q-functions, actor-free approach.
result Performance and sample efficiency comparable to actor-critic methods.
New criteria for non-isometric group actions in metric spaces.
problem Understanding group actions on non-isometric spaces.
method Generalizing results from isometric to continuous group actions.
result Criterion for cocompact cyclic groups to be inessential.
RANDPOL uses randomized networks for efficient reinforcement learning in continuous state and action MDPs.
problem Efficient reinforcement learning in environments with continuous state and action spaces.
method RANDPOL uses randomized function approximation to represent policy and value functions, providing finite time guarantees and improved numerical performance.
result RANDPOL achieves better numerical performance and provides finite time guarantees compared to deep neural network based algorithms.
We develop a method for optimizing policies with continuous actions using observational data.
problem Optimizing policies with continuous actions using observational data where the data collection policy is unknown.
method A semi-parametric approach with a doubly robust off-policy estimate.
result Our method is robust to estimation errors of the policy function or the regression model.
ZoomRL learns efficient strategies for large state-action spaces using a metric.
problem Handling large state-action spaces in reinforcement learning.
method ZoomRL leverages continuous bandits to adaptively discretize the joint space.
result Achieves worst-case regret of $ ilde{O}(H^{rac{5}{2}} K^{rac{d+1}{d+2}})$.
Unified family of estimators for bounded action spaces reduces variance in policy gradients.
problem High variance in policy gradients for bounded action spaces.
method Marginal policy gradients family of estimators for directional control.
result APG estimator offers substantial improvement over standard policy gradient.
DiGrad improves multi-task reinforcement learning in robotic systems.
problem Efficient multi-task reinforcement learning in complex robotic systems with shared actions.
method Differential Policy Gradient (DiGrad) for simultaneous training of multiple tasks in a single actor-critic network.
result DiGrad outperforms related methods in continuous action spaces, supporting efficient multi-task learning.
A hybrid model combines Q-learning and PID controller for continuous vehicle control.
problem Learning unsatisfactory results with discrete action space in autonomous driving.
method Combining Q-learning and PID controller, using Quadratic Q-function approximation and action network.
result Autonomous vehicle successfully learns smooth and efficient driving behavior.
SDPG algorithm improves sample efficiency and reward in DRL for continuous action spaces.
problem Improving sample efficiency and reward in distributional reinforcement learning for continuous action spaces.
method SDPG algorithm models return distribution using samples via reparameterization technique.
result SDPG shows better sample efficiency and higher reward in OpenAI Gym environments.
New findings on when to use action space exploration in reinforcement learning.
problem Understanding when to use action space exploration over traditional methods.
method Theoretical analysis and empirical testing of simple exploration methods.
result Exploration in action space is preferred when parametric complexity exceeds action space dimensionality and horizon length.
Connectedness proved for Zd actions on 1D manifolds by C2 diffeomorphisms.
problem Connectedness of Zd actions by C2 diffeomorphisms on 1D manifolds. method Proved connectedness through continuous paths of C1+ac diffeomorphisms. result Connectedness of Zd actions by C2 diffeomorphisms on 1D manifolds. Efficient algorithm for reinforcement learning in large state-action spaces with adaptive discretization.
problem Efficient reinforcement learning in large, potentially continuous state-action spaces.
method Adaptive Q-learning policy with data-driven adaptive discretization. result Demonstrates improved performance compared to existing methods, especially in adapting to the problem's structure.
Study shows conditions for continuity of foliated homeomorphisms action on space of leaves.
problem Conditions for continuity of foliated homeomorphisms action on space of leaves.
method Investigated sufficient conditions for continuity of the homomorphism ψ: H(X, Δ) → H(Y) induced by the action of foliated homeomorphisms on the space of leaves.
result Similar results hold for a more general class of partitions of locally compact Hausdorff spaces.
New analysis shows black-box methods outperform action space methods in certain scenarios.
problem Comparing black-box methods vs. action space methods in exploration.
method Theoretical analyses and empirical comparisons of simple methods on various problems.
result Complexity of exploration in parameter space depends on parameter space dimensionality, while action space complexity depends on both action space and horizon length.
New algorithm learns optimal actions in complex decision problems.
problem Optimal policy learning in continuous Markov decision problems.
method Nonparametric stochastic compositional gradient descent in RKHS.
result Algorithm converges to optimal policies with low Bellman error.
New method improves policy gradient performance in continuous control tasks.
problem Improving policy gradient methods for continuous control tasks.
method Numerical integration approach to all-action policy gradient.
result Improved performance and sample efficiency in continuous control tasks.
Study compares SPG and PPO for racing games, finding SPG more stable with weighted actions.
problem Training continuous action reinforcement learning algorithms for racing games.
method Introduced novel racing environment, tested SPG and PPO with modifications and experience replay.
result Experience replay not beneficial for PPO in continuous action spaces, SPG more stable with weighted actions.
A new method for continuous control avoids local movement issues.
problem Limitations of policy gradient methods in continuous control.
method Distributional framework and Generative Actor Critic (GAC) method.
result GAC outperforms policy gradient methods in continuous domains.
Paper solves POMDPs in continuous time and discrete spaces.
problem Optimal decision making in discrete state and action space systems under partial observability.
method Combining optimal filtering theory and deep learning to solve a Hamilton-Jacobi-Bellman equation.
result Derives a mathematical description and solution approach for continuous-time POMDPs.
Enhances reinforcement learning safety through risk-averse exploration.
problem Safety concerns in reinforcement learning due to sub-optimal actions.
method Distributionally robust policy iteration scheme with lower bound guarantees.
result Efficient algorithm that prevents poor decisions and converges to optimal policy.
New framework discovers non-affine continuous symmetries in neural networks.
problem Lack of efficient methods for detecting non-affine continuous symmetries in neural networks.
method Computational framework for discovering infinitesimal generators of multi-parameter group actions.
result Framework can discover non-affine continuous symmetries in neural networks.
New algorithm for context bandits with continuous actions.
problem Efficient decision-making with unknown action structures.
method Reduction-style algorithm combining supervised learning.
result Proven to work in general and validated with experiments.