A new ML algorithm solves complex economic control problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
Study optimal consumption with drawdown limits over a fixed time frame.
Paper tackles utility maximization with job-switching and retirement constraints.
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
Microgrids (MGs) are small, local power grids that can operate independently from the larger utility grid. Combined with the Internet of Things (IoT), a smart MG can leverage the sensory data and machine learning techniques for intelligent energy management. This paper focuses on deep reinforcement learning (DRL)-based…
This paper concerns the numerical solution of the finite-horizon Optimal Investment problem with transaction costs under Potential Utility. The problem is initially posed in terms of an evolutive HJB equation with gradient constraints. In Finite-Horizon Optimal Investment with Transaction Costs: A Parabolic Double Obst…
A new method optimizes spatial sampling for level set estimation in one dimension.
New algorithms minimize regret in SSP with optimal sparse updates.
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
Paper solves portfolio problem using improved stochastic methods.
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
In this expository paper we illustrate the generality of game theoretic probability protocols of Shafer and Vovk (2001) in finite-horizon discrete games. By restricting ourselves to finite-horizon discrete games, we can explicitly describe how discrete distributions with finite support and the discrete pricing formulas…
UCB-Advantage learns MDPs with regret.
We study a stochastic, continuous time model on a finite horizon for a firm that produces a single good. We model the production capacity as an Ito diffusion controlled by a nondecreasing process representing the cumulative investment. The firm aims to maximize its expected total net profit by choosing the optimal inve…
Consider the problem of sampling sequentially from a finite number of populations, specified by random variables , and ; where denotes the outcome from population the time it is sampled. It is assumed that for each fixed , $\{ X^i_k \}_{k …
This paper investigates the problem of maximizing expected terminal utility in a discrete-time financial market model with a finite horizon under non-dominated model uncertainty. We use a dynamic programming framework together with measurable selection arguments to prove that under mild integrability conditions, an opt…
We characterise the value function of the optimal dividend problem with a finite time horizon as the unique classical solution of a suitable Hamilton-Jacobi-Bellman equation. The optimal dividend strategy is realised by a Skorokhod reflection of the fund's value at a time-dependent optimal boundary. Our results are obt…
Motivated by recent axiomatic developments, we study the risk- and ambiguity-averse investment problem where trading takes place over a fixed finite horizon and terminal payoffs are evaluated according to a criterion defined in terms of a quasiconcave utility functional. We extend to the present setting certain existen…
In this research we study a finite horizon optimal purchasing problem for items with a mean reverting price process. Under this model a fixed amount of identical items are bought under a given deadline, with the objective of minimizing the cost of their purchasing price and associated holding cost. We prove that the op…
I analyse the frequentist regret of the famous Gittins index strategy for multi-armed bandits with Gaussian noise and a finite horizon. Remarkably it turns out that this approach leads to finite-time regret guarantees comparable to those available for the popular UCB algorithm. Along the way I derive finite-time bounds…
Solves portfolio optimization with costs using numerical methods.
We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoints. Analogous to TS…
Optimizes latency and false alarm probability in change detection problems.
We study a stochastic control approach to managed futures portfolios. Building on the Schwartz 97 stochastic convenience yield model for commodity prices, we formulate a utility maximization problem for dynamically trading a single-maturity futures or multiple futures contracts over a finite horizon. By analyzing the a…
New meta-reinforcement learning method improves performance in finite-horizon MDPs.
New algorithm reduces regret in CMDPs without cancellation of errors.
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.
We explicitly solve the nonlinear PDE that is the continuous limit of dynamic programming of \emph{expert prediction problem} in finite horizon setting with experts. The \emph{expert prediction problem} is formulated as a zero sum game between a player and an adversary. By showing that the solution is $\mathcal{C…
This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …
We prove a general duality result for multi-stage portfolio optimization problems in markets with proportional transaction costs. The financial market is described by Kabanov's model of foreign exchange markets over a finite probability space and finite-horizon discrete time steps. This framework allows us to compare v…
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
We consider classical Merton problem of terminal wealth maximization in finite horizon. We assume that the drift of the stock is following Ornstein-Uhlenbeck process and the volatility of it is following GARCH(1) process. In particular, both mean and volatility are unbounded. We assume that there is Knightian uncertain…
Algorithm reduces episode count for CMDPs with constraints.
We study the portfolio selection problem of a long-run investor who is maximising the asymptotic growth rate of her expected utility. We show that, somewhat surprisingly, it is essentially not affected by introduction of a floor constraint which requires the wealth process to dominate a given benchmark at all times. We…
Deep Galerkin Method estimates value function for mean-field control problem.
In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed dynamic optimal control and stopping problems in the existing literature, to stu…
New algorithm reduces offline RL data requirements significantly.
Solves optimal stopping for Gauss-Markov bridges using time-space transformation.
Deep learning solves complex stochastic control with jumps.
Optimal policy for multi-armed multi-action bandits with unknown parameters.
This paper concerns the numerical solution of a fully nonlinear parabolic double obstacle problem arising from a finite portfolio selection with proportional transaction costs. We consider the optimal allocation of wealth among multiple stocks and a bank account in order to maximize the finite horizon discounted utilit…
Q-MMR evaluates policies using reweighted rewards and moment matching.
Study long-term asset liquidation behavior with external flows.
This paper presents several numerical applications of deep learning-based algorithms that have been introduced in [HPBL18]. Numerical and comparative tests using TensorFlow illustrate the performance of our different algorithms, namely control learning by performance iteration (algorithms NNcontPI and ClassifPI), contr…
Derives time-averaged active inference from control principles.
We study several optimal stopping problems that arise from trading a mean-reverting price spread over a finite horizon. Modeling the spread by the Ornstein-Uhlenbeck process, we analyze three different trading strategies: (i) the long-short strategy; (ii) the short-long strategy, and (iii) the chooser strategy, i.e. th…