Overparameterized models generalize well in offline contextual bandits, but policy-based algorithms struggle.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Step-DAD improves BED by periodically updating a design policy during experiments.
This work improves policy-based training by proposing an evaluation balance objective for GFlowNets.
This work explores representation complexity in RL paradigms, revealing model-based RL as the easiest task.
New GFlowNet training framework using policy gradients for combinatorial object generation.
Modern deep learning methods provide effective means to learn good representations. However, is a good representation itself sufficient for sample efficient reinforcement learning? This question has largely been studied only with respect to (worst-case) approximation error, in the more classical approximate dynamic pro…
New algorithm reduces dynamic regret for MDPs with unknown transition and adversarial rewards.
New method defends RL agents from poisoning attacks without MDP knowledge.
Develops ODRPO to improve RL algorithms with better performance and stability.
We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization. Specifically, we show that softmax consistent action values correspond to optimal entropy regularized policy pro…
The study optimizes free trial lengths to boost subscriptions and consumer loyalty.
The goal of reinforcement learning (RL) is to let an agent learn an optimal control policy in an unknown environment so that future expected rewards are maximized. The model-free RL approach directly learns the policy based on data samples. Although using many samples tends to improve the accuracy of policy learning, c…
MAVEN improves multi-agent exploration by hybridizing value and policy-based methods.
Optimizing option exercise policies based on variance optimal martingale measure can lead to unappealing results.
New algorithm reduces RL policy optimization gap.
We solve POMDPs by approximating them as finite-state MDPs.
Deep Reinforcement Learning (DRL) has become a powerful strategy to solve complex decision making problems based on Deep Neural Networks (DNNs). However, it is highly data demanding, so unfeasible in physical systems for most applications. In this work, we approach an alternative Interactive Machine Learning (IML) stra…
Survey of RL in finance, tackling complex decision-making.
New RL algorithm tackles non-stationary environments with flexible policy updates.
Study on adversarial training's impact on deep neural reinforcement learning policies.
A new imitation learning method uses random search for simple policies, outperforming complex models.
Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority without testing the new policy? To answer this question, we introduce the G-SCOPE algorithm that evaluates a new policy based on data generated by…
This work highlights problems with off-policy estimation in recommender systems due to unobserved confounders.
As the Portable Document Format (PDF) file format increases in popularity, research in analysing its structure for text extraction and analysis is necessary. Detecting headings can be a crucial component of classifying and extracting meaningful data. This research involves training a supervised learning model to detect…
In decision making problems for continuous state and action spaces, linear dynamical models are widely employed. Specifically, policies for stochastic linear systems subject to quadratic cost functions capture a large number of applications in reinforcement learning. Selected randomized policies have been studied in th…
New algorithm finds near-optimal policies efficiently in zero-sum games.
Privileged Information Dropout improves RL performance without distillation.
We introduce Bayesian least-squares policy iteration (BLSPI), an off-policy, model-free, policy iteration algorithm that uses the Bayesian least-squares temporal-difference (BLSTD) learning algorithm to evaluate policies. An online variant of BLSPI has been also proposed, called randomised BLSPI (RBLSPI), that improves…
In artificial intelligence, we often specify tasks through a reward function. While this works well in some settings, many tasks are hard to specify this way. In deep reinforcement learning, for example, directly specifying a reward as a function of a high-dimensional observation is challenging. Instead, we present an …
Learning goal-oriented dialogues by means of deep reinforcement learning has recently become a popular research topic. However, commonly used policy-based dialogue agents often end up focusing on simple utterances and suboptimal policies. To mitigate this problem, we propose a class of novel temperature-based extension…
Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical success, its underlying mathematical principle on {\em policy-distri…
Data augmentation is commonly used to encode invariances in learning methods. However, this process is often performed in an inefficient manner, as artificial examples are created by applying a number of transformations to all points in the training set. The resulting explosion of the dataset size can be an issue in te…
In this paper, making use of recent statistical physics techniques and models, we address the specific role of randomness in financial markets, both at the micro and the macro level. In particular, we review some recent results obtained about the effectiveness of random strategies of investment, compared with some of t…
Dual behavior policy improves reinforcement learning across various environments.
Gradient-EM Bayesian meta-learning accelerates adaptation with reduced computation and improved robustness.
Policy analysts wish to visualize a range of policies for large simulator-defined Markov Decision Processes (MDPs). One visualization approach is to invoke the simulator to generate on-policy trajectories and then visualize those trajectories. When the simulator is expensive, this is not practical, and some method is r…
We study the problem of off-policy value evaluation in reinforcement learning (RL), where one aims to estimate the value of a new policy based on data collected by a different policy. This problem is often a critical step when applying RL in real-world problems. Despite its importance, existing general methods either h…
Imitation learning is an effective alternative approach to learn a policy when the reward function is sparse. In this paper, we consider a challenging setting where an agent and an expert use different actions from each other. We assume that the agent has access to a sparse reward function and state-only expert observa…
In this paper, we study a multi-step interactive recommendation problem, where the item recommended at current step may affect the quality of future recommendations. To address the problem, we develop a novel and effective approach, named CFRL, which seamlessly integrates the ideas of both collaborative filtering (CF) …
The paper reclassifies RL algorithms using inference concepts.
A network of spiking agents learns complex tasks using global reward signals.
New Bellman error estimator improves offline model selection performance.
Active learning method improves local model validity estimation.
iDAD uses neural networks to quickly adapt experiments without likelihoods.
This paper improves MARL for networked systems through new protocols and discount factors.
We show that compact fully connected (FC) deep learning networks trained to classify wireless protocols using a hierarchy of multiple denoising autoencoders (AEs) outperform reference FC networks trained in a typical way, i.e., with a stochastic gradient based optimization of a given FC architecture. Not only is the co…
There are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems. In this setting, an online user is the environment; neither the reward function nor the environment dynamics are clearly defined, making the application of RL challenging. In this paper, we propose a …
VA-OPE improves OPE by incorporating variance information, achieving tighter error bounds.