We investigate the growth optimal strategy over a finite time horizon for a stock and bond portfolio in an analytically solvable multiplicative Markovian market model. We show that the optimal strategy consists in holding the amount of capital invested in stocks within an interval around an ideal optimal investment. Th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study examines Wang-Yau quasi-local energy in strong fields near apparent horizons.
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
A new ML algorithm solves complex economic control problems.
We consider generic static spacetimes with Killing horizons and study properties of curvature tensors in the horizon limit. It is determined that the Weyl, Ricci, Riemann and Einstein tensors are algebraically special and mutually aligned on the horizon. It is also pointed out that results obtained in the tetrad adjust…
Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.
In this expository paper we illustrate the generality of game theoretic probability protocols of Shafer and Vovk (2001) in finite-horizon discrete games. By restricting ourselves to finite-horizon discrete games, we can explicitly describe how discrete distributions with finite support and the discrete pricing formulas…
This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …
New algorithms minimize regret in SSP with optimal sparse updates.
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
Modeling risk and performance with Levy-stable distributions.
I analyse the frequentist regret of the famous Gittins index strategy for multi-armed bandits with Gaussian noise and a finite horizon. Remarkably it turns out that this approach leads to finite-time regret guarantees comparable to those available for the popular UCB algorithm. Along the way I derive finite-time bounds…
In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed dynamic optimal control and stopping problems in the existing literature, to stu…
Study optimal consumption with drawdown limits over a fixed time frame.
Microgrids (MGs) are small, local power grids that can operate independently from the larger utility grid. Combined with the Internet of Things (IoT), a smart MG can leverage the sensory data and machine learning techniques for intelligent energy management. This paper focuses on deep reinforcement learning (DRL)-based…
Study examines deformations of Kerr-(A)dS near horizon geometry.
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes. We analyze the variance of the most popular approaches through the viewpoint of con…
New method detects black hole horizons using Lie algebra invariants.
Study long-term asset liquidation behavior with external flows.
Firms miscount their customers who stop buying without saying goodbye.
Developed LQ MFG theory with common noise, proving existence and uniqueness.
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
We characterise the value function of the optimal dividend problem with a finite time horizon as the unique classical solution of a suitable Hamilton-Jacobi-Bellman equation. The optimal dividend strategy is realised by a Skorokhod reflection of the fund's value at a time-dependent optimal boundary. Our results are obt…
Paper analyzes convergence of dynamic policy gradient for MDPs, improving performance in finite-time problems.
Optimal reinsurance and dividend strategy for insurance companies in a finite time.
Logarithmic regret achieved in continuous-time linear-quadratic reinforcement learning.
Study on BSDEs with random time horizon, focusing on existence and properties.
This paper is concerned with offline reinforcement learning (RL), which learns using pre-collected data without further exploration. Effective offline RL would be able to accommodate distribution shift and limited data coverage. However, prior algorithms or analyses either suffer from suboptimal sample complexities or …
In reinforcement learning, the discount factor controls the agent's effective planning horizon. Traditionally, this parameter was considered part of the MDP; however, as deep reinforcement learning algorithms tend to become unstable when the effective planning horizon is long, recent works refer to as a hyper-p…
We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoints. Analogous to TS…
The paper studies the question of whether the classical mirror and synchronous couplings of two Brownian motions minimise and maximise, respectively, the coupling time of the corresponding geometric Brownian motions. We establish a characterisation of the optimality of the two couplings over any finite time horizon and…
In this paper, we propose to combine imitation and reinforcement learning via the idea of reward shaping using an oracle. We study the effectiveness of the near-optimal cost-to-go oracle on the planning horizon and demonstrate that the cost-to-go oracle shortens the learner's planning horizon as function of its accurac…
We demonstrate the existence of spherically-symmetric truly naked black holes (TNBH) for which the Kretschmann scalar is finite on the horizon but some curvature components including those responsible for tidal forces as well as the energy density measured by a free-falling observer are infinite. We choose a ra…
This paper concerns the numerical solution of the finite-horizon Optimal Investment problem with transaction costs under Potential Utility. The problem is initially posed in terms of an evolutive HJB equation with gradient constraints. In Finite-Horizon Optimal Investment with Transaction Costs: A Parabolic Double Obst…
New algorithm learns optimal policies with just 1 episode, settling horizon-dependence in RL.
Develops a regression approach for solving MDPs with general state and action spaces.
New meta-reinforcement learning method improves performance in finite-horizon MDPs.
Optimizes latency and false alarm probability in change detection problems.
Algorithm reduces episode count for CMDPs with constraints.
Study optimal liquidation strategies with infinite horizon and regime switching.
Paper tackles utility maximization with job-switching and retirement constraints.
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
We consider the problem of portfolio optimization in a simple incomplete market and under a general utility function. By working with the associated Hamilton-Jacobi-Bellman partial differential equation (HJB PDE), we obtain a closed-form formula for a trading strategy which approximates the optimal trading strategy whe…
This paper improves Thompson Sampling for complex decision-making problems.
New algorithm reduces offline RL data requirements significantly.
I introduce and analyse an anytime version of the Optimally Confident UCB (OCUCB) algorithm designed for minimising the cumulative regret in finite-armed stochastic bandits with subgaussian noise. The new algorithm is simple, intuitive (in hindsight) and comes with the strongest finite-time regret guarantees for a hori…