New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes. We analyze the variance of the most popular approaches through the viewpoint of con…
Firms miscount their customers who stop buying without saying goodbye.
This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …
A new method optimizes spatial sampling for level set estimation in one dimension.
Logarithmic regret achieved in continuous-time linear-quadratic reinforcement learning.
New method estimates off-policy data without needing known behavior policy.
Kernel-UCBVI algorithm balances exploration and exploitation in metric state-action spaces.
We introduce wavelet-based methodology for estimation of realized variance allowing its measurement in the time-frequency domain. Using smooth wavelets and Maximum Overlap Discrete Wavelet Transform, we allow for the decomposition of the realized variance into several investment horizons and jumps. Basing our estimator…
New meta-reinforcement learning method improves performance in finite-horizon MDPs.
Develops a regression approach for solving MDPs with general state and action spaces.
Optimal reinsurance and dividend strategy for insurance companies in a finite time.
A new Bayesian method optimizes time-dependent expensive functions with lookahead.
Paper solves portfolio problem using improved stochastic methods.
We investigate the growth optimal strategy over a finite time horizon for a stock and bond portfolio in an analytically solvable multiplicative Markovian market model. We show that the optimal strategy consists in holding the amount of capital invested in stocks within an interval around an ideal optimal investment. Th…
We consider the problem of off-policy evaluation for reinforcement learning, where the goal is to estimate the expected reward of a target policy using offline data collected by running a logging policy . Standard importance-sampling based approaches for this problem suffer from a variance that scales exponentia…
Study examines Wang-Yau quasi-local energy in strong fields near apparent horizons.
We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
A new ML algorithm solves complex economic control problems.
Anticipatory portfolios use richer models to optimize investments.
We consider generic static spacetimes with Killing horizons and study properties of curvature tensors in the horizon limit. It is determined that the Weyl, Ricci, Riemann and Einstein tensors are algebraically special and mutually aligned on the horizon. It is also pointed out that results obtained in the tetrad adjust…
We study the online estimation of the optimal policy of a Markov decision process (MDP). We propose a class of Stochastic Primal-Dual (SPD) methods which exploit the inherent minimax duality of Bellman equations. The SPD methods update a few coordinates of the value and policy estimates as a new state transition is obs…
Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.
In this expository paper we illustrate the generality of game theoretic probability protocols of Shafer and Vovk (2001) in finite-horizon discrete games. By restricting ourselves to finite-horizon discrete games, we can explicitly describe how discrete distributions with finite support and the discrete pricing formulas…
New algorithms minimize regret in SSP with optimal sparse updates.
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG algorithm suffers from high sample complexity. In this paper we consider the deterministic value gradients to improve the sample efficiency of …
We develop methods to approximate derivatives for causal inference problems using data.
The study shows black hole horizons at low temperatures have limited topology.
Modeling risk and performance with Levy-stable distributions.
I analyse the frequentist regret of the famous Gittins index strategy for multi-armed bandits with Gaussian noise and a finite horizon. Remarkably it turns out that this approach leads to finite-time regret guarantees comparable to those available for the popular UCB algorithm. Along the way I derive finite-time bounds…
In this paper, we investigate dynamic optimization problems featuring both stochastic control and optimal stopping in a finite time horizon. The paper aims to develop new methodologies, which are significantly different from those of mixed dynamic optimal control and stopping problems in the existing literature, to stu…
In this paper, we analyze the finite sample complexity of stochastic system identification using modern tools from machine learning and statistics. An unknown discrete-time linear system evolves over time under Gaussian noise without external inputs. The objective is to recover the system parameters as well as the Kalm…
Proposes pT-Learning for optimal dynamic treatment regimes in mHealth.
Estimates roughness of financial volatility paths using horizontal visibility graphs.
Study optimal consumption with drawdown limits over a fixed time frame.
Microgrids (MGs) are small, local power grids that can operate independently from the larger utility grid. Combined with the Internet of Things (IoT), a smart MG can leverage the sensory data and machine learning techniques for intelligent energy management. This paper focuses on deep reinforcement learning (DRL)-based…
Paper studies apparent horizon dynamics and introduces a null comparison principle.
An optimal algorithm for multi-armed bandits with constraints.
Study examines deformations of Kerr-(A)dS near horizon geometry.
We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…
New method detects black hole horizons using Lie algebra invariants.
Deep Galerkin Method estimates value function for mean-field control problem.
Study long-term asset liquidation behavior with external flows.
In this global study of solutions to the linear wave equation on Schwarzschild de Sitter spacetimes we attend to the cosmological region of spacetime which is bounded in the past by cosmological horizons and to the future by a spacelike hypersurface at infinity. We prove an energy estimate capturing the expansion of th…
We present a continuous-time maximum likelihood estimation methodology for credit rating transition probabilities, taking into account the presence of censored data. We perform rolling estimates of the transition matrices with exponential time weighting with varying horizons and discuss the underlying dynamics of trans…
This paper studies a recent proposal to use randomized value functions to drive exploration in reinforcement learning. These randomized value functions are generated by injecting random noise into the training data, making the approach compatible with many popular methods for estimating parameterized value functions. B…