We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study optimal liquidation strategies with infinite horizon and regime switching.
Solves infinite horizon portfolio problem with path-dependent labor income.
Paper proposes an efficient online learning method using an offline dataset for infinite horizon MDPs.
Investment and consumption strategy for risk-averse agents with Epstein-Zin utility.
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estimating the stationar…
Extends utility maximization theory for infinite horizons without strong no-arbitrage assumptions.
In an incomplete market, with incompleteness stemming from stochastic factors imperfectly correlated with the underlying stocks, we derive representations of homothetic (power, exponential and logarithmic) forward performance processes in factor-form using ergodic BSDE. We also develop a connection between the forward …
Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduced for learning infinite-horizon average-reward Markov Decision Processes (MDPs). The first algorithm reduces the problem to the discounted-r…
A new model-free algorithm achieves near-optimal regret for infinite-horizon MDPs.
New algorithms for learning MDPs with linear approximations in infinite-horizon settings.
Develops a method to estimate policy values robustly in the presence of confounding variables.
Study tackles OPE in confounded settings, estimating policy value from proxies.
UCRL2-VTR achieves nearly optimal regret for learning MDPs with linear function approximation.
New algorithm reduces reinforcement learning regret to sqrt(T) without strong dynamics assumptions.
Study optimal portfolio management with periodic evaluations in stochastic models, considering convex constraints.
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
Study optimal consumption and investment for investors with Epstein-Zin preferences.
We estimate risk measures in Markov cost processes with lower and upper bounds.
Derives time-averaged active inference from control principles.
This paper improves Thompson Sampling for complex decision-making problems.
We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite horizon context for off-policy policy evaluation. We propose estimated mixture policy …
New algorithm reduces regret in infinite MDPs with optimal variance-dependent bounds.
A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-learning with…
The application of existing methods for constructing optimal dynamic treatment regimes is limited to cases where investigators are interested in optimizing a utility function over a fixed period of time (finite horizon). In this manuscript, we develop an inferential procedure based on temporal difference residuals for …
The paper analyzes the sample complexity of offline RL with linear approximations, identifying a hard regime and providing an algorithm.
New algorithm LOOP learns infinite-horizon AMDPs efficiently with function approximation.
We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…
We solve non-Markovian optimal switching problems in discrete time on an infinite horizon, when the decision maker is risk aware and the filtration is general, and establish existence and uniqueness of solutions for the associated reflected backward stochastic difference equations. An example application to hydropower …
Algorithm converges to Nash equilibria in competitive games.
This paper optimizes portfolio management in incomplete markets with stochastic factors, considering periodic wealth evaluations.
New algorithms learn MDPs with better regret bounds using generative sampling.
In this paper, we examine higher order difference problems. Using the "squeezing" argument, we derive both Euler's condition and the transversality condition. In order to derive the two conditions, two needed assumptions are identified. A counterexample, in which the transversality condition is not satisfied without th…
Improved RL algorithm with linear MDPs for offline learning with partial data coverage.
Gaussian processes provide a flexible framework for forecasting, removing noise, and interpreting long temporal datasets. State space modelling (Kalman filtering) enables these non-parametric models to be deployed on long datasets by reducing the complexity to linear in the number of data points. The complexity is stil…
We discuss a class of risk-sensitive portfolio optimization problems. We consider the portfolio optimization model investigated by Nagai in 2003. The model by its nature can include fixed income securities as well in the portfolio. Under fairly general conditions, we prove the existence of optimal portfolio in both fin…
New algorithms reduce dynamic regret in online MDPs with changing losses.
Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG algorithm suffers from high sample complexity. In this paper we consider the deterministic value gradients to improve the sample efficiency of …
New method estimates off-policy data without needing known behavior policy.
Actor-Critic method achieves optimal regret for unichain MDPs.
This paper considers a mortgage contract where the borrower pays a fixed mortgage rate and has the choice of making prepayment. Assume the market interest follows the CIR model, a free boundary problem is formulated. Here we focus on the infinite horizon problem. Using variational method, we obtain an analytical soluti…
The classical optimal investment and consumption problem with infinite horizon is studied in the presence of transaction costs. Both proportional and fixed costs as well as general utility functions are considered. Weak dynamic programming is proved in the general setting and a comparison result for possibly discontinu…
The article constructs a forward utility for markets with multiple default risks.
We consider the problem of maximizing expected power utility from consumption over an infinite horizon in the Black-Scholes model with proportional transaction costs, as studied in Shreve and Soner [Ann. Appl. Probab. 4 (1994) 609-692]. Similar to Kallsen and Muhle-Karbe [Ann. Appl. Probab. 20 (2010) 1341-1358], we der…
We consider an impulse control problem in infinite horizon applied with switching technology. We suppose that the firm decides at certain moments (impulse moments) to switch technology, leading to a jump of the firm value. We show that the value function for such problems satisfies a dynamic programming principle versi…
This is a brief technical note to clarify some of the issues with applying the application of the algorithm posterior sampling for reinforcement learning (PSRL) in environments without fixed episodes. In particular, this paper aims to: - Review some of results which have been proven for finite horizon MDPs (Osband et a…
Reinforcement learning is a general technique that allows an agent to learn an optimal policy and interact with an environment in sequential decision making problems. The goodness of a policy is measured by its value function starting from some initial state. The focus of this paper is to construct confidence intervals…
Proposes pT-Learning for optimal dynamic treatment regimes in mHealth.