The paper improves energy decay estimates for Dir-stationary Q-valued functions and applies them to Liouville-type theorems and continuity.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Q-Distribution Guided Q-Learning corrects overestimation of uncertain OOD actions in offline RL.
Investigates Q value evolution in Stable Baselines for DQL in simple vs complex environments.
Proves continuity and singular set dimension for 2D maps with Q values.
A new DQN algorithm improves portfolio management and risk assessment in digital assets.
We investigate statistical uncertainty quantification for reinforcement learning (RL) and its implications in exploration policy. Despite ever-growing literature on RL applications, fundamental questions about inference and error quantification, such as large-sample behaviors, appear to remain quite open. In this paper…
A new method stabilizes deep reinforcement learning by using QGraphs to retain replay memory information.
Proposes Optimistic Pessimistically Initialised Q-Learning (OPIQ) for better exploration in RL.
The paper proves the convergence of Q-value for Gaussian rewards.
We analyze a notion of multiple valued sections of a vector bundle over an abstract smooth Riemannian manifold, which was suggested by W. Allard in the unpublished note "Some useful techniques for dealing with multiple valued functions" and generalizes Almgren's -valued functions. We study some relevant properties o…
New approach optimizes policies in adversarial MDPs using adversarial learning.
New method finds unseen states for RL, improving performance.
We methodologically address the problem of Q-value overestimation in deep reinforcement learning to handle high-dimensional state spaces efficiently. By adapting concepts from information theory, we introduce an intrinsic penalty signal encouraging reduced Q-value estimates. The resultant algorithm encompasses a wide r…
Paper improves RL algorithms with known optimal Q-value predictions.
Proposes a value-based method for continuous control without an actor.
Improved Q-learning for multi-agent reinforcement learning by weighting joint action values.
KL regularization helps RL algorithms by implicitly averaging q-values.
We give a formula of the Upsilon invariant of any L-space cable knot using and . The integral value of the Upsilon invariant gives a -valued knot concordance invariant. We compute the integral values for L-space iterated cable knots.
The paper proposes a method to infer Q-values online with Q-Learning.
Study Q-learning with averaging for reinforcement learning, proving efficient inference and error bounds.
Being able to reach any desired location in the environment can be a valuable asset for an agent. Learning a policy to navigate between all pairs of states individually is often not feasible. An all-goals updating algorithm uses each transition to learn Q-values towards all goals simultaneously and off-policy. However …
Policy gradient is an efficient technique for improving a policy in a reinforcement learning setting. However, vanilla online variants are on-policy only and not able to take advantage of off-policy data. In this paper we describe a new technique that combines policy gradient with off-policy Q-learning, drawing experie…
Network slicing promises to provision diversified services with distinct requirements in one infrastructure. Deep reinforcement learning (e.g., deep -learning, DQL) is assumed to be an appropriate algorithm to solve the demand-aware inter-slice resource management issue in network slicing by regarding the …
Bayesian approach improves -greedy exploration in RL.
New method estimates optimal Q-values with better accuracy for specific problems.
We construct Lipschitz -valued functions which approximate carefully integral currents when their cylindrical excess is small and they are almost minimizing in a suitable sense. This result is used in two subsequent works to prove the discreteness of the singular set for the following three classes of -dimensiona…
In a discounted reward Markov Decision Process (MDP), the objective is to find the optimal value function, i.e., the value function corresponding to an optimal policy. This problem reduces to solving a functional equation known as the Bellman equation and a fixed point iteration scheme known as the value iteration is u…
We study reinforcement learning (RL) in high dimensional episodic Markov decision processes (MDP). We consider value-based RL when the optimal Q-value is a linear function of d-dimensional state-action feature representation. For instance, in deep-Q networks (DQN), the Q-value is a linear function of the feature repres…
We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance p…
Bayes-UCBVI tackles reinforcement learning with a new upper confidence bound method.
In this paper we study the singular set of Dirichlet-minimizing -valued maps from into a smooth compact manifold without boundary. Similarly to what happens in the case of single valued minimizing harmonic maps, we show that this set is always -rectifiable with uniform Minkowski b…
Reinforcement learning agents are faced with two types of uncertainty. Epistemic uncertainty stems from limited data and is useful for exploration, whereas aleatoric uncertainty arises from stochastic environments and must be accounted for in risk-sensitive applications. We highlight the challenges involved in simultan…
A RL approach optimizes metal AM process parameters for consistent melt pool depth.
Instability and variability of Deep Reinforcement Learning (DRL) algorithms tend to adversely affect their performance. Averaged-DQN is a simple extension to the DQN algorithm, based on averaging previously learned Q-values estimates, which leads to a more stable training procedure and improved performance by reducing …
The paper proposes a structure learning model for efficient reinforcement learning.
Partially observable Markov decision processes (POMDPs) with continuous state and observation spaces have powerful flexibility for representing real-world decision and control problems but are notoriously difficult to solve. Recent online sampling-based algorithms that use observation likelihood weighting have shown un…
We develop a multivalued theory for the stability operator of (a constant multiple of) a minimally immersed submanifold of a Riemannian manifold . We define the multiple valued counterpart of the classical Jacobi fields as the minimizers of the second variation functional defined on a Sobolev space of …
Improves reinforcement learning extrapolation in Gridworlds.
Q-Ensembles are a model-free approach where input images are fed into different Q-networks and exploration is driven by the assumption that uncertainty is proportional to the variance of the output Q-values obtained. They have been shown to perform relatively well compared to other exploration strategies. Further, mode…
Paper provides estimates for varifolds with critical mean curvature.
Study invariants of elliptic curves in LCS manifolds, leading to new phenomena in Riemann-Finsler geometry.
New RL theory predicts deep RL success based on greedy actions under random policies.
Equivariant CNNs improve RL performance in symmetric environments.
Deep Reinforcement Learning (RL) recently emerged as one of the most competitive approaches for learning in sequential decision making problems with fully observable environments, e.g., computer Go. However, very little work has been done in deep RL to handle partially observable environments. We propose a new architec…
We present a method for a certain class of Markov Decision Processes (MDPs) that can relate the optimal policy back to one or more reward sources in the environment. For a given initial state, without fully computing the value function, q-value function, or the optimal policy the algorithm can determine which rewards w…
Explains agent behavior through intended outcomes in reinforcement learning.
Optimizes mobile notifications for multiple objectives using reinforcement learning.
Optimal trade execution is an important problem faced by essentially all traders. Much research into optimal execution uses stringent model assumptions and applies continuous time stochastic control to solve them. Here, we instead take a model free approach and develop a variation of Deep Q-Learning to estimate the opt…