Stochastic Q-learning tackles large action spaces with reduced computation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified reinforcement learning and stochastic processes with action-driven processes.
The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available decisions (actions) at each time step is stochastic. Recently, the stochastic action set Markov decision process (SAS-MDP) formulation has b…
Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value gradients is desirable as policy improvement occurs along the direction of stee…
Quantum stochastic flow computes heat kernel traces for Ricci flat manifolds.
The paper tackles counterfactual learning for stochastic policies with continuous actions.
Method predicts future rewards from past actions in a linear Gaussian system.
Algorithm maximizes rewards with a budget and giving up option.
Method trains neural network for optimal decisions from stochastic simulators.
FinFlowRL combines imitation and reinforcement learning for better financial control.
Model-free learning for multi-agent stochastic games is an active area of research. Existing reinforcement learning algorithms, however, are often restricted to zero-sum games, and are applicable only in small state-action spaces or other simplified settings. Here, we develop a new data efficient Deep-Q-learning method…
New algorithm reduces regret in private online learning with optimal gap-dependent rate.
TOFU-POV tackles partially observed linear bandits, achieving sublinear regret with low-dimensional action vectors.
New algorithm reduces sleeping bandits' regret to O(sqrt(T)).
New attack manipulates UCB algorithm, new defense algorithm reduces pseudo-regret.
Diffusion models mimic human actions in sequential tasks.
Investigates spontaneous symmetry breaking in non-equilibrium systems.
Paper solves stochastic contextual linear bandits using linear bandit algorithms.
We studied isometric stochastic flows of a Stratonovich stochastic differential equation on spheres, i.e. on the standard sphere and Gromoll-Meyer exotic sphere. The standard sphere can be constructed as the quotient manifold with the so-called -action of , where…
Study shows how to learn optimal policies quickly in stochastic control problems.
New algorithms handle unpredictable actions in sequential learning.
New method for linear bandits with unknown sparsity, improving sparse regret bounds.
PQR estimates reward functions from actions and states without assuming state-only rewards.
FinFlowRL learns from experts to optimize financial control in changing markets.
Continuous reinforcement learning such as DDPG and A3C are widely used in robot control and autonomous driving. However, both methods have theoretical weaknesses. While DDPG cannot control noises in the control process, A3C does not satisfy the continuity conditions under the Gaussian policy. To address these concerns,…
This paper investigates the adversarial Bandits with Knapsack (BwK) online learning problem, where a player repeatedly chooses to perform an action, pays the corresponding cost, and receives a reward associated with the action. The player is constrained by the maximum budget that can be spent to perform actions, an…
Improved GFlowNets learn more efficiently with trajectory balance.
The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…
New approach handles stochastic and partially-observable environments using discrete autoencoders and Monte Carlo tree search.
User engagement in social networks depends critically on the number of online actions their users take in the network. Can we design an algorithm that finds when to incentivize users to take actions to maximize the overall activity in a social network? In this paper, we model the number of online actions over time usin…
We consider stochastic multi-armed bandit problems with graph feedback, where the decision maker is allowed to observe the neighboring actions of the chosen action. We allow the graph structure to vary with time and consider both deterministic and Erdős-Rényi random graph models. For such a graph feedback model, we fir…
Stochastic games provide a framework for interactions among multiple agents and enable a myriad of applications. In these games, agents decide on actions simultaneously, the state of every agent moves to the next state, and each agent receives a reward. However, finding an equilibrium (if exists) in this game is often …
Develops a machine learning framework for computing most probable paths in stochastic systems.
Improved reinforcement learning for episodes with varying action sets.
New method uses hindsight to make exploration robust in stochastic environments.
We consider Markov Decision Problems defined over continuous state and action spaces, where an autonomous agent seeks to learn a map from its states to actions so as to maximize its long-term discounted accumulation of rewards. We address this problem by considering Bellman's optimality equation defined over action-val…
User engagement in online social networking depends critically on the level of social activity in the corresponding platform--the number of online actions, such as posts, shares or replies, taken by their users. Can we design data-driven algorithms to increase social activity? At a user level, such algorithms may incre…
Study robust control for systems with continuous states using adversarial perturbations.
In many practical problems, a learning agent may want to learn the best action in hindsight without ever taking a bad action, which is significantly worse than the default production action. In general, this is impossible because the agent has to explore unknown actions, some of which can be bad, to learn better action…
We propose expected policy gradients (EPG), which unify stochastic policy gradients (SPG) and deterministic policy gradients (DPG) for reinforcement learning. Inspired by expected sarsa, EPG integrates (or sums) across actions when estimating the gradient, instead of relying only on the action in the sampled trajectory…
Improved ExO method achieves near-optimal bounds in both stochastic and adversarial settings.
We explore Deep Reinforcement Learning in a parameterized action space. Specifically, we investigate how to achieve sample-efficient end-to-end training in these tasks. We propose a new compact architecture for the tasks where the parameter policy is conditioned on the output of the discrete action policy. We also prop…
New method reduces ensemble size for linear bandits, achieving near optimal regret.
New algorithms reduce regret in both stochastic and adversarial partial monitoring problems.
Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…
We derive a consistent differential representation for the dynamics of a self-financing portfolio for different hedging strategies. In the basis of the derivation there is the so called "retarded action principle", which represents the causality in the evolution of dependent stochastic variables. We demonstrate this pr…
We study adversarial attacks that manipulate the reward signals to control the actions chosen by a stochastic multi-armed bandit algorithm. We propose the first attack against two popular bandit algorithms: -greedy and UCB, \emph{without} knowledge of the mean rewards. The attacker is able to spend only logarithmic …
We propose expected policy gradients (EPG), which unify stochastic policy gradients (SPG) and deterministic policy gradients (DPG) for reinforcement learning. Inspired by expected sarsa, EPG integrates across the action when estimating the gradient, instead of relying only on the action in the sampled trajectory. We es…