Proposes a method to learn policies from offline data with reduced bias.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Introduces LTQL for factored policies in cooperative MARL.
FACMAC combines deep policy gradients with factored critic for multi-agent reinforcement learning.
This paper studies optimal approximation factors in misspecified off-policy RL, identifying key factors under various settings.
The study examines how global economic policy uncertainty affects crude oil futures volatility.
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning and generalization by explicitly conditioning the policy on a parametric context space. In this paper…
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different "contexts". Bayesian optimization approaches to contextual policy search (CPS) offer data-efficient policy learning that generalize over a context space. We propose to impr…
Paper extends quantile factor analysis with probabilistic methods for better economic policy and financial condition prediction.
RL learns to ignore factors in factor investing portfolios.
Policy gradient methods with aggregated states can achieve better performance than approximate policy iteration.
The paper solves multi-period portfolio selection with constraints using a dynamic factor model.
Study finds relevance of exchange and inflation rates to economic factors.
N-discount optimality was introduced as a hierarchical form of policy- and value-function optimality, with Blackwell optimality lying at the top level of the hierarchy Veinott (1969); Blackwell (1962). We formalize notions of myopic discount factors, value functions and policies in terms of Blackwell optimality in MDPs…
FPGs use structure to improve policy learning in complex tasks.
Policy analysts wish to visualize a range of policies for large simulator-defined Markov Decision Processes (MDPs). One visualization approach is to invoke the simulator to generate on-policy trajectories and then visualize those trajectories. When the simulator is expensive, this is not practical, and some method is r…
Deep Reinforcement Learning (DRL) algorithms for continuous action spaces are known to be brittle toward hyperparameters as well as \cut{being}sample inefficient. Soft Actor Critic (SAC) proposes an off-policy deep actor critic algorithm within the maximum entropy RL framework which offers greater stability and empiric…
We propose a novel approach to train a multi-modal policy from mixed demonstrations without their behavior labels. We develop a method to discover the latent factors of variation in the demonstrations. Specifically, our method is based on the variational autoencoder with a categorical latent variable. The encoder infer…
We derive an optimal policy for adaptively restarting a randomized algorithm, based on observed features of the run-so-far, so as to minimize the expected time required for the algorithm to successfully terminate. Given a suitable Bayesian prior, this result can be used to select the optimal black-box optimization algo…
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficien…
Paper proposes a new REINFORCE algorithm for mining formulaic alpha factors with reduced variance.
New method selects best offline RL policies from logged data.
Paper tackles offline RL with weak assumptions on both function classes and data coverage.
In this paper we study a model-based approach to calculating approximately optimal policies in Markovian Decision Processes. In particular, we derive novel bounds on the loss of using a policy derived from a factored linear model, a class of models which generalize numerous previous models out of those that come with s…
Study minimax off-policy evaluation in multi-armed bandits with known and unknown behavior policies.
Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority without testing the new policy? To answer this question, we introduce the G-SCOPE algorithm that evaluates a new policy based on data generated by…
Study classifies liability insurance policies using machine learning.
It has been postulated that a good representation is one that disentangles the underlying explanatory factors of variation. However, it remains an open question what kind of training framework could potentially achieve that. Whereas most previous work focuses on the static setting (e.g., with images), we postulate that…
Study shows sample complexity for learning optimal policies in SSP with generative model.
The paper tackles robust policy learning in MDPs using statistical methods.
Method selects best estimator for off-policy evaluation.
This paper proposes non-stationary factor models for financial stress in the UK.
The paper studies how neural policies can be interpreted using decision trees.
It has been postulated that a good representation is one that disentangles the underlying explanatory factors of variation. However, it remains an open question what kind of training framework could potentially achieve that. Whereas most previous work focuses on the static setting (e.g., with images), we postulate that…
New method speeds up lifelong learning of complex tasks.
Data-driven RL solves Merton's expected utility problem via policy randomization.
Improved model-free RL algorithm with reduced sample complexity.
We study a reinforcement learning setting, where the state transition function is a convex combination of a stochastic continuous function and a deterministic function. Such a setting generalizes the widely-studied stochastic state transition setting, namely the setting of deterministic policy gradient (DPG). We firstl…
Recently, \citet{SuttonMW15} introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD involves a contraction operator, with a -contraction modulus (where is the …
Current economic theories miss most of economic dynamics.
CPPO learns policies from partial offline data in MDPs with structural assumptions.
We consider a large, homogeneous portfolio of life or disability annuity policies. The policies are assumed to be independent conditional on an external stochastic process representing the economic-demographic environment. Using a conditional law of large numbers, we establish the connection between claims reserving an…
We investigate the impact of Knightian uncertainty on the optimal timing policy of an ambiguity averse decision maker in the case where the underlying factor dynamics follow a multidimensional Brownian motion and the exercise payoff depends on either a linear combination of the factors or the radial part of the driving…
New method reduces bias and variance in OPE for large action spaces.
Adaptive sequential decision making is one of the central challenges in machine learning and artificial intelligence. In such problems, the goal is to design an interactive policy that plans for an action to take, from a finite set of actions, given some partial observations. It has been shown that in many applicat…
New Q-learning algorithm reduces sample complexity for large discount factors.
This paper compares two stock factor models in China's A-share market.
Develops a statistical model for SOFR term structure in incomplete markets.
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based search to improve policies online. During search, the simulation policy is adapted to explore the most promising lines of play. MCTS has been used by state-of-the-art programs for many problems, however a disadvantage to MCTS is that it estimates the va…