A new ML algorithm solves complex economic control problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithms minimize regret in SSP with optimal sparse updates.
Study optimal consumption with drawdown limits over a fixed time frame.
In this expository paper we illustrate the generality of game theoretic probability protocols of Shafer and Vovk (2001) in finite-horizon discrete games. By restricting ourselves to finite-horizon discrete games, we can explicitly describe how discrete distributions with finite support and the discrete pricing formulas…
Paper uses DRL for smart MG energy dispatch, improving stability and performance.
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
I analyse the frequentist regret of the famous Gittins index strategy for multi-armed bandits with Gaussian noise and a finite horizon. Remarkably it turns out that this approach leads to finite-time regret guarantees comparable to those available for the popular UCB algorithm. Along the way I derive finite-time bounds…
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
Paper tackles utility maximization with job-switching and retirement constraints.
New IS methods fail to reduce variance in long-horizon MDPs.
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
A new method optimizes spatial sampling for level set estimation in one dimension.
New meta-reinforcement learning method improves performance in finite-horizon MDPs.
We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.
Solves expert prediction problem for 4 experts in finite time horizon.
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
UCB-Advantage learns MDPs with regret.
This paper concerns the numerical solution of the finite-horizon Optimal Investment problem with transaction costs under Potential Utility. The problem is initially posed in terms of an evolutive HJB equation with gradient constraints. In Finite-Horizon Optimal Investment with Transaction Costs: A Parabolic Double Obst…
New algorithms ensure policies perform at least as good as a baseline in reinforcement learning.
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
Improved RL algorithm reduces regret in large state spaces.
Paper solves portfolio problem using improved stochastic methods.
Q-MMR evaluates policies using reweighted rewards and moment matching.
In this research we study a finite horizon optimal purchasing problem for items with a mean reverting price process. Under this model a fixed amount of identical items are bought under a given deadline, with the objective of minimizing the cost of their purchasing price and associated holding cost. We prove that the op…
We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
This paper investigates the problem of maximizing expected terminal utility in a discrete-time financial market model with a finite horizon under non-dominated model uncertainty. We use a dynamic programming framework together with measurable selection arguments to prove that under mild integrability conditions, an opt…
New algorithm reduces regret in CMDPs without cancellation of errors.
Consider the problem of sampling sequentially from a finite number of populations, specified by random variables , and ; where denotes the outcome from population the time it is sampled. It is assumed that for each fixed , $\{ X^i_k \}_{k …
We study a stochastic, continuous time model on a finite horizon for a firm that produces a single good. We model the production capacity as an Ito diffusion controlled by a nondecreasing process representing the cumulative investment. The firm aims to maximize its expected total net profit by choosing the optimal inve…
Firms miscount their customers who stop buying without saying goodbye.
Optimizes latency and false alarm probability in change detection problems.
Algorithm reduces episode count for CMDPs with constraints.
New model selects robustly in adversarial reinforcement learning with unknown corruption.
This paper studies a recent proposal to use randomized value functions to drive exploration in reinforcement learning. These randomized value functions are generated by injecting random noise into the training data, making the approach compatible with many popular methods for estimating parameterized value functions. B…
This paper studies the properties of the optimal portfolio-consumption strategies in a {finite horizon} robust utility maximization framework with different borrowing and lending rates. In particular, we allow for constraints on both investment and consumption strategies, and model uncertainty on both drift and volatil…
Motivated by recent axiomatic developments, we study the risk- and ambiguity-averse investment problem where trading takes place over a fixed finite horizon and terminal payoffs are evaluated according to a criterion defined in terms of a quasiconcave utility functional. We extend to the present setting certain existen…
New algorithm reduces offline RL data requirements significantly.
We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoints. Analogous to TS…
This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …
New RL difficulty shown for discounted settings.
We study a stochastic control approach to managed futures portfolios. Building on the Schwartz 97 stochastic convenience yield model for commodity prices, we formulate a utility maximization problem for dynamically trading a single-maturity futures or multiple futures contracts over a finite horizon. By analyzing the a…
In this paper, we obtain the finite-horizon and infinite-horizon ruin probability asymptotics for risk processes with claims of subexponential tails for non-stationary arrival processes that satisfy a large deviation principle. As a result, the arrival process can be dependent, non-stationary and non-renewal. We give t…
This paper is concerned with offline reinforcement learning (RL), which learns using pre-collected data without further exploration. Effective offline RL would be able to accommodate distribution shift and limited data coverage. However, prior algorithms or analyses either suffer from suboptimal sample complexities or …
We characterise the value function of the optimal dividend problem with a finite time horizon as the unique classical solution of a suitable Hamilton-Jacobi-Bellman equation. The optimal dividend strategy is realised by a Skorokhod reflection of the fund's value at a time-dependent optimal boundary. Our results are obt…
We study the portfolio selection problem of a long-run investor who is maximising the asymptotic growth rate of her expected utility. We show that, somewhat surprisingly, it is essentially not affected by introduction of a floor constraint which requires the wealth process to dominate a given benchmark at all times. We…
Portfolio turnpikes state that, as the investment horizon increases, optimal portfolios for generic utilities converge to those of isoelastic utilities. This paper proves three kinds of turnpikes. In a general semimartingale setting, the abstract turnpike states that optimal final payoffs and portfolios converge under …
Solves portfolio optimization with costs using numerical methods.
Solves optimal stopping for Gauss-Markov bridges using time-space transformation.