Q-Distribution Guided Q-Learning corrects overestimation of uncertain OOD actions in offline RL.
problem Overestimation of Q-values for out-of-distribution actions in offline reinforcement learning.
method QDQ applies a pessimistic adjustment to Q-values in uncertain OOD regions based on a consistency model.
result QDQ improves performance on the D4RL benchmark and achieves significant improvements across many tasks.
New method finds unseen states for RL, improving performance.
problem Offline RL struggles with unseen states and actions.
method Value-informed state perturbations and filtering.
result Improved performance in offline RL tasks.
CQL learns conservative Q-functions to improve offline RL performance.
problem Leveraging large, static datasets in reinforcement learning without further interaction.
method Conservative Q-learning (CQL) which learns a conservative Q-function to lower-bound policy values.
result CQL substantially outperforms existing offline RL methods, often achieving 2-5 times higher final returns.
The paper improves energy decay estimates for Dir-stationary Q-valued functions and applies them to Liouville-type theorems and continuity.
problem Improving energy decay estimates for Dir-stationary Q-valued functions.
method Establishing improved decay estimates and applying them to derive Liouville-type theorems and continuity.
result Dir-stationary Q-valued functions exhibit the Lebesgue property and reside in a generalized Campanato-Morrey space.
Model-based reinforcement learning algorithms tend to achieve higher sample efficiency than model-free methods. However, due to the inevitable errors of learned models, model-based methods struggle to achieve the same asymptotic performance as model-free methods. In this paper, We propose a Policy Optimization method w…
Investigates Q value evolution in Stable Baselines for DQL in simple vs complex environments.
problem DQL in Stable Baselines struggles with simple non-game environments.
method Comparison of TrafficLight and FrozenLake environments; Q value decomposition analysis.
result Q values meander far from optimal in complex relationships between states.
Proves continuity and singular set dimension for 2D maps with Q values.
problem Interior regularity of 2D Q-valued maps. method Strong concentration-compactness theorem for equicontinuous maps.
result 2D Q-valued maps are Hölder continuous with singular set dimension ≤1. Optimizes mobile notifications for multiple objectives using reinforcement learning.
problem Optimizing mobile notification systems for multiple objectives.
method End-to-end offline reinforcement learning with Double Deep Q-network and Conservative Q-learning.
result Demonstrates improved performance and benefits of the proposed approach.
A new DQN algorithm improves portfolio management and risk assessment in digital assets.
problem Singular prediction mode and limited data source in deep learning models for asset management.
method Introduced DQN algorithm into asset management portfolios, considering market risk.
result Performance exceeds benchmark, proving DRL algorithm's effectiveness in portfolio management.
We investigate statistical uncertainty quantification for reinforcement learning (RL) and its implications in exploration policy. Despite ever-growing literature on RL applications, fundamental questions about inference and error quantification, such as large-sample behaviors, appear to remain quite open. In this paper…
A new method stabilizes deep reinforcement learning by using QGraphs to retain replay memory information.
problem Stabilizing model-free off-policy deep reinforcement learning with soft divergence.
method Representing past experiences as a QGraph, selecting a subgraph with favorable structure, and using lower bounds for temporal difference learning.
result QG-DDPG method is less prone to soft divergence and more robust to hyperparameters.
Proposes Optimistic Pessimistically Initialised Q-Learning (OPIQ) for better exploration in RL.
problem Pessimistic initialisation of Q-values in deep RL leads to poor exploration performance.
method Augments pessimistically initialised Q-values with count-based bonuses to ensure optimism.
result OPIQ outperforms non-optimistic DQN variants in hard exploration tasks.
The paper proves the convergence of Q-value for Gaussian rewards.
problem Existing proofs cannot guarantee convergence of the Q-function for Gaussian rewards.
method Using the central limit theorem and relaxing the condition to E[r(s,a)2]<∞. result Proves the convergence of the Q-function under the condition of E[r(s,a)2]<∞. We analyze a notion of multiple valued sections of a vector bundle over an abstract smooth Riemannian manifold, which was suggested by W. Allard in the unpublished note "Some useful techniques for dealing with multiple valued functions" and generalizes Almgren's Q-valued functions. We study some relevant properties o…
New approach optimizes policies in adversarial MDPs using adversarial learning.
problem Optimizing policies in adversarial Markov decision processes.
method Adversarial learning on advantage functions, extending previous reductions.
result Stronger regret criteria and performance guarantees for policy optimization.
We methodologically address the problem of Q-value overestimation in deep reinforcement learning to handle high-dimensional state spaces efficiently. By adapting concepts from information theory, we introduce an intrinsic penalty signal encouraging reduced Q-value estimates. The resultant algorithm encompasses a wide r…
Paper improves RL algorithms with known optimal Q-value predictions.
problem Improving RL algorithms with known optimal Q-value predictions.
method Shows how to improve regret bounds by distilling known optimal Q-value predictions.
result Achieves sublinear regret with known predictions, even if not distillation.
Proposes a value-based method for continuous control without an actor.
problem Computational infeasibility of evaluating Q-values in continuous action spaces.
method Structurally maximizable Q-functions, actor-free approach.
result Performance and sample efficiency comparable to actor-critic methods.
Improved Q-learning for multi-agent reinforcement learning by weighting joint action values.
problem QMIX restricts Q-values to monotonic mixtures, limiting complex value functions. method Introduced weighted projection to recover optimal policies, improving performance.
result CW QMIX and OW QMIX outperform baseline QMIX on multi-agent tasks.
KL regularization helps RL algorithms by implicitly averaging q-values.
problem Understanding why KL regularization improves RL performance.
method An approximate value iteration scheme, studying KL and entropy regularization.
result Strong performance bound combining linear horizon dependency and averaging effect of estimation errors.
We give a formula of the Upsilon invariant of any L-space cable knot Kp,q using p,ΥK and ΥTp,q. The integral value of the Upsilon invariant gives a Q-valued knot concordance invariant. We compute the integral values for L-space iterated cable knots.
Theoretical analysis confirms non-conservative algorithms can converge to optimal policies.
problem Theoretical guarantees for non-conservative reinforcement learning algorithms.
method Theoretical analysis of Peng's Q(λ) algorithm. result Peng's Q(λ) converges to an optimal policy under certain conditions. Study finds conserved quantities for two types of curves on conformal sphere.
problem Identifying conserved quantities for specific types of curves on a conformal sphere.
method Used parallel tractor and Lagrangian formalism to compute conserved quantities.
result Found relation between conserved quantities of two curve types.
The paper proposes a method to infer Q-values online with Q-Learning.
problem High variance and instability in reinforcement learning algorithms.
method Adapting FCLT for a modified Q-learning approach and constructing confidence intervals.
result The proposed method provides more stable and reliable inference of Q-values.
Survey on conservation laws for geometric PDEs.
problem Modeling polyharmonic maps.
method Conservation law approach.
result Overview of conservation laws in geometric PDEs.
The article discusses conservation laws for polyharmonic maps and their applications.
problem Understanding conservation laws for polyharmonic maps.
method Recalling the stress-energy tensor and showing conservation laws with Killing vector fields.
result Conservation laws for polyharmonic maps and their applications.
Study Q-learning with averaging for reinforcement learning, proving efficient inference and error bounds.
problem Efficient inference and error bounds for Q-learning with averaging.
method Functional central limit theorem and asymptotic linear estimator for optimal Q-value function.
result Standardized partial-sum process converges weakly to a rescaled Brownian motion, matching instance-dependent lower bound for error.
Paper presents a reduction-based framework for conservative bandits and RL with improved lower and upper bounds.
problem Conservative bandits and reinforcement learning problems.
method Reduction technique to calculate necessary and sufficient budget from baseline policy.
result Improved lower and upper bounds for various conservative settings.
I consider the existence and structure of conservation laws for the general class of evolutionary scalar second-order differential equations with parabolic symbol. First I calculate the linearized characteristic cohomology for such equations. This provides an auxiliary differential equation satisfied by the conservatio…
Being able to reach any desired location in the environment can be a valuable asset for an agent. Learning a policy to navigate between all pairs of states individually is often not feasible. An all-goals updating algorithm uses each transition to learn Q-values towards all goals simultaneously and off-policy. However …
Policy gradient is an efficient technique for improving a policy in a reinforcement learning setting. However, vanilla online variants are on-policy only and not able to take advantage of off-policy data. In this paper we describe a new technique that combines policy gradient with off-policy Q-learning, drawing experie…
Non-trivial conservation law found for a specific system.
problem Conservation law for a specific system with a vanishing characteristic.
method Analyzing overdetermined system with given characteristics.
result Non-trivial conservation law despite vanishing characteristic.
This work connects symmetries and conserved quantities in machine learning.
problem Improving machine learning models by learning conserved quantities.
method Using Noether's theorem, learn symmetries and conserved quantities directly from data.
result Correctly identifies conserved quantities and improves model performance.
Proposes a conservative exploration method for RL agents.
problem Guaranteeing performance of exploratory policies in RL.
method Importance sampling for off-policy policy evaluation.
result Derives a regret bound ensuring no conservative constraint violation.
Network slicing promises to provision diversified services with distinct requirements in one infrastructure. Deep reinforcement learning (e.g., deep Q-learning, DQL) is assumed to be an appropriate algorithm to solve the demand-aware inter-slice resource management issue in network slicing by regarding the …
Conservation law for weakly harmonic mappings in high dimensions.
problem Conservation law for harmonic mappings in supercritical dimensions.
method Partial extension of Rivière's conservation law with Lorentz integrability condition.
result Conservation law for weakly harmonic mappings in supercritical dimensions.
The conservation laws of the third order quasilinear scalar evolution equations are considered via differential system and characteristic cohomology. We find a subspace of 2 forms in the infinite prolonged space in which every conservation law has a unique representative. The structure of this subspace naturally gives …
Bayesian approach improves ε-greedy exploration in RL.
problem Improving ε-greedy exploration in model-free RL. method Introducing a Bayesian model update for ε based on BMC. result Proposed ε- exttt{BMC} algorithm efficiently balances exploration and exploitation. Given a vector field on a manifold M, we define a globally conserved quantity to be a differential form whose Lie derivative is exact. Integrals of conserved quantities over suitable submanifolds are constant under time evolution, the Kelvin circulation theorem being a well-known special case. More generally, conserved…
New method estimates optimal Q-values with better accuracy for specific problems.
problem Estimating optimal Q-values in reinforcement learning is difficult and varies by problem instance.
method Local minimax framework and variance-reduced Q-learning.
result Sharp lower bounds on estimation accuracy for Q-learning.
We study higher-order conservation laws of the non-linearizable elliptic Poisson equation ∂z∂zˉ∂2u=−f(u) as elements of the characteristic cohomology of the associated exterior differential system. The theory of characteristic cohomology determines a normal form for diffe…
New conservation laws found for polyharmonic maps in critical dimension.
problem Existence of conservation laws for polyharmonic maps in critical dimension.
method Small perturbation of Uhlenbeck's gauge fixing matrix.
result Existence of conservation laws for elliptic systems of even order in critical dimension.
A challenging problem in complex networks is the network reconstruction problem from data. This work deals with a class of networks denoted as conserved networks, in which a flow associated with every edge and the flows are conserved at all non-source and non-sink nodes. We propose a novel polynomial time algorithm to …
MC-LSTM extends LSTM to conserve mass in neural networks.
problem Conservation laws in real-world systems.
method Extending LSTM's inductive bias to conserve mass.
result MC-LSTM sets new state-of-the-art for predicting peak flows.
The paper studies symmetries and conservation laws of non-diagonalisable hydrodynamic systems.
problem Integrating non-diagonalisable hydrodynamic systems of partial differential equations.
method Analysis of gl-regular Nijenhuis operators, splitting Theorem for symmetries and conservation laws, relationship between symmetries and conservation laws.
result The system of partial differential equations is integrable in quadratures.
We present a connection between the Killing fields that arise in the loop-group approach to integrable systems and conservation laws viewed as elements of the characteristic cohomology. We use the connection to generate the complete set of conservation laws (as elements of the characteristic cohomology) for the Tzitzei…
A new algorithm balances exploration and exploitation in online decision-making.
problem Balancing exploration and exploitation in online decision-making.
method Proposed C4-UCB algorithm incorporating conservative mechanism. result Proved n-step upper regret bound for two situations.
We obtain necessary and sufficient conditions for the existence of "conservation laws" on null hypersurfaces for the wave equation on general four-dimensional Lorentzian manifolds. Examples of null hypersurfaces exhibiting such conservation laws include the standard null cones of Minkowski spacetime and the degenerate …