We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study optimal liquidation strategies with infinite horizon and regime switching.
Improved algorithm for optimal stopping problems reduces runtime.
In an incomplete market, with incompleteness stemming from stochastic factors imperfectly correlated with the underlying stocks, we derive representations of homothetic (power, exponential and logarithmic) forward performance processes in factor-form using ergodic BSDE. We also develop a connection between the forward …
Paper proposes an efficient online learning method using an offline dataset for infinite horizon MDPs.
Extends utility maximization theory for infinite horizons without strong no-arbitrage assumptions.
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
Investment and consumption strategy for risk-averse agents with Epstein-Zin utility.
Developed LQ MFG theory with common noise, proving existence and uniqueness.
Deep neural nets approximate random dynamical system trajectories uniformly in time.
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estimating the stationar…
We consider an infinite horizon portfolio problem with borrowing constraints, in which an agent receives labor income which adjusts to financial market shocks in a path dependent way. This path-dependency is the novelty of the model, and leads to an infinite dimensional stochastic optimal control problem. We solve the …
Study tackles OPE in confounded settings, estimating policy value from proxies.
UCRL2-VTR achieves nearly optimal regret for learning MDPs with linear function approximation.
Paper proves existence of ambient manifolds for null hypersurfaces solving Einstein equations.
Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduced for learning infinite-horizon average-reward Markov Decision Processes (MDPs). The first algorithm reduces the problem to the discounted-r…
Derives time-averaged active inference from control principles.
A new model-free algorithm achieves near-optimal regret for infinite-horizon MDPs.
We present a new infinite class of near-horizon geometries of degenerate horizons, satisfying Einstein's equations for all odd dimensions greater than five. The symmetry and topology of these solutions is compatible with those of black holes. The simplest examples give horizons of spatial topology S^3xS^2 or the non-tr…
We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…
Firms miscount their customers who stop buying without saying goodbye.
New algorithms for learning MDPs with linear approximations in infinite-horizon settings.
New insights into black hole horizons from asymptotic expansions.
A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-learning with…
Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG algorithm suffers from high sample complexity. In this paper we consider the deterministic value gradients to improve the sample efficiency of …
We demonstrate the existence of spherically-symmetric truly naked black holes (TNBH) for which the Kretschmann scalar is finite on the horizon but some curvature components including those responsible for tidal forces as well as the energy density measured by a free-falling observer are infinite. We choose a ra…
This paper improves Thompson Sampling for complex decision-making problems.
The application of existing methods for constructing optimal dynamic treatment regimes is limited to cases where investigators are interested in optimizing a utility function over a fixed period of time (finite horizon). In this manuscript, we develop an inferential procedure based on temporal difference residuals for …
New method estimates off-policy data without needing known behavior policy.
This paper is concerned with offline reinforcement learning (RL), which learns using pre-collected data without further exploration. Effective offline RL would be able to accommodate distribution shift and limited data coverage. However, prior algorithms or analyses either suffer from suboptimal sample complexities or …
New algorithm reduces reinforcement learning regret to sqrt(T) without strong dynamics assumptions.
In this paper, we provide an elementary, unified treatment of two distinct blue-shift instabilities for the scalar wave equation on a fixed Kerr black hole background: the celebrated blue-shift at the Cauchy horizon (familiar from the strong cosmic censorship conjecture) and the time-reversed red-shift at the event hor…
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
New algorithm reduces regret in infinite MDPs with optimal variance-dependent bounds.
Study optimal portfolio management with periodic evaluations in stochastic models, considering convex constraints.
New algorithms learn MDPs with better regret bounds using generative sampling.
Study optimal consumption and investment for investors with Epstein-Zin preferences.
We estimate risk measures in Markov cost processes with lower and upper bounds.
We prove that a large class of smooth solutions to the linear wave equation on subextremal rotating Kerr spacetimes which are regular and decaying along the event horizon become singular at the Cauchy horizon. More precisely, we show that assuming appropriate upper and lower bounds on the energy along t…
Proposes pT-Learning for optimal dynamic treatment regimes in mHealth.
We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite horizon context for off-policy policy evaluation. We propose estimated mixture policy …
The paper analyzes the sample complexity of offline RL with linear approximations, identifying a hard regime and providing an algorithm.
New algorithms reduce dynamic regret in online MDPs with changing losses.
The infinite Viterbi alignment is the limiting maximum a-posteriori estimate of the unobserved path in a hidden Markov model as the length of the time horizon grows. For models on state-space satisfying a new ``decay-convexity'' condition, we develop an approach to existence of the infinite Viterbi ali…
Overview of risk-sensitive Markov decision processes with Optimized Certainty Equivalent.
We prove that any smooth vacuum spacetime containing a compact Cauchy horizon with surface gravity that can be normalised to a non-zero constant admits a Killing vector field. This proves a conjecture by Moncrief and Isenberg from 1983 under the assumption on the surface gravity and generalises previous results due to …
In this paper, we obtain the finite-horizon and infinite-horizon ruin probability asymptotics for risk processes with claims of subexponential tails for non-stationary arrival processes that satisfy a large deviation principle. As a result, the arrival process can be dependent, non-stationary and non-renewal. We give t…
Study long-term asset liquidation behavior with external flows.