New algorithms minimize regret in SSP with optimal sparse updates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Q-MMR evaluates policies using reweighted rewards and moment matching.
We develop a dual-control method for approximating investment strategies in incomplete environments that emerge from the presence of trading constraints. Convex duality enables the approximate technology to generate lower and upper bounds on the optimal value function. The mechanism rests on closed-form expressions per…
A large proportion of market making models derive from the seminal model of Avellaneda and Stoikov. The numerical approximation of the value function and the optimal quotes in these models remains a challenge when the number of assets is large. In this article, we propose closed-form approximations for the value functi…
A new ML algorithm solves complex economic control problems.
Control of non-episodic, finite-horizon dynamical systems with uncertain dynamics poses a tough and elementary case of the exploration-exploitation trade-off. Bayesian reinforcement learning, reasoning about the effect of actions and future observations, offers a principled solution, but is intractable. We review, then…
Study optimal consumption with drawdown limits over a fixed time frame.
Microgrids (MGs) are small, local power grids that can operate independently from the larger utility grid. Combined with the Internet of Things (IoT), a smart MG can leverage the sensory data and machine learning techniques for intelligent energy management. This paper focuses on deep reinforcement learning (DRL)-based…
In this expository paper we illustrate the generality of game theoretic probability protocols of Shafer and Vovk (2001) in finite-horizon discrete games. By restricting ourselves to finite-horizon discrete games, we can explicitly describe how discrete distributions with finite support and the discrete pricing formulas…
We study the finite horizon Merton portfolio optimization problem in a general local-stochastic volatility setting. Using model coefficient expansion techniques, we derive approximations for the both the value function and the optimal investment strategy. We also analyze the `implied Sharpe ratio' and derive a series a…
We consider the exploration-exploitation dilemma in finite-horizon reinforcement learning (RL). When the state space is large or continuous, traditional tabular approaches are unfeasible and some form of function approximation is mandatory. In this paper, we introduce an optimistically-initialized variant of the popula…
Deep Galerkin Method estimates value function for mean-field control problem.
Algorithm reduces episode count for CMDPs with constraints.
This paper develops algorithms for high-dimensional stochastic control problems based on deep learning and dynamic programming. Unlike classical approximate dynamic programming approaches, we first approximate the optimal policy by means of neural networks in the spirit of deep reinforcement learning, and then the valu…
New model selects robustly in adversarial reinforcement learning with unknown corruption.
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
Finite-horizon sequential experimental design (SED) arises naturally in many contexts, including hyperparameter tuning in machine learning among more traditional settings. Computing the optimal policy for such problems requires solving Bellman equations, which are generally intractable. Most existing work resorts to se…
I analyse the frequentist regret of the famous Gittins index strategy for multi-armed bandits with Gaussian noise and a finite horizon. Remarkably it turns out that this approach leads to finite-time regret guarantees comparable to those available for the popular UCB algorithm. Along the way I derive finite-time bounds…
Deep learning solves complex stochastic control with jumps.
This paper is about index policies for minimizing (frequentist) regret in a stochastic multi-armed bandit model, inspired by a Bayesian view on the problem. Our main contribution is to prove that the Bayes-UCB algorithm, which relies on quantiles of posterior distributions, is asymptotically optimal when the reward dis…
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
Paper tackles utility maximization with job-switching and retirement constraints.
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
A new method optimizes spatial sampling for level set estimation in one dimension.
A new method for risk-averse decision-making in Markov processes with improved regret bounds.
New meta-reinforcement learning method improves performance in finite-horizon MDPs.
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.
This paper concerns the numerical solution of a fully nonlinear parabolic double obstacle problem arising from a finite portfolio selection with proportional transaction costs. We consider the optimal allocation of wealth among multiple stocks and a bank account in order to maximize the finite horizon discounted utilit…
VPR improves posterior uncertainty quantification by combining VI and predictive resampling.
In this paper a quantitative analysis of the ruin probability in finite time of discrete risk process with proportional reinsurance and investment of finance surplus is focused on. It is assumed that the total loss on a unit interval has a light-tailed distribution -- exponential distribution and a heavy-tailed distrib…
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
UCB-Advantage learns MDPs with regret.
We study the problem of optimal trading using general alpha predictors with linear costs and temporary impact. We do this within the framework of stochastic optimization with finite horizon using both limit and market orders. Consistently with other studies, we find that the presence of linear costs induces a no-tradin…
Develops a new risk measure for Markov chains' asymptotic behavior.
NVMDP framework tackles non-stationary MDPs with varying discount rates.
We present an approach for pricing European call options in presence of proportional transaction costs, when the stock price follows a general exponential Lévy process. The model is a generalization of the celebrated work of Davis, Panas and Zariphopoulou (1993), where the value of the option is defined as the utility …
This paper concerns the numerical solution of the finite-horizon Optimal Investment problem with transaction costs under Potential Utility. The problem is initially posed in terms of an evolutive HJB equation with gradient constraints. In Finite-Horizon Optimal Investment with Transaction Costs: A Parabolic Double Obst…
Study on fake stationary Volterra Heston model for non-stationary processes.
New algorithms learn MDPs with better regret bounds using generative sampling.
We consider a finite horizon optimal stopping problem related to trade-off strategies between expected profit and cost cash-flows of an investment under uncertainty. The optimal problem is first formulated in terms of a system of Snell envelopes for the profit and cost yields which act as obstacles to each other. We th…
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
Paper solves portfolio problem using improved stochastic methods.
In this research we study a finite horizon optimal purchasing problem for items with a mean reverting price process. Under this model a fixed amount of identical items are bought under a given deadline, with the objective of minimizing the cost of their purchasing price and associated holding cost. We prove that the op…
Our understanding of reinforcement learning (RL) has been shaped by theoretical and empirical results that were obtained decades ago using tabular representations and linear function approximators. These results suggest that RL methods that use temporal differencing (TD) are superior to direct Monte Carlo estimation (M…
We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
New method efficiently evaluates policies using trajectory data.
We study online reinforcement learning for finite-horizon deterministic control systems with {\it arbitrary} state and action spaces. Suppose that the transition dynamics and reward function is unknown, but the state and action space is endowed with a metric that characterizes the proximity between different states and…
Improved stochastic approximation method reduces residual error.