UCRL-WVTR tackles long-term reinforcement learning with general approximations, achieving horizon-free and instance-dependent regret bounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study optimal liquidation strategies with infinite horizon and regime switching.
This paper studies the utility maximization problem with changing time horizons in the incomplete Brownian setting. We first show that the primal value function and the optimal terminal wealth are continuous with respect to the time horizon . Secondly, we exemplify that the expected utility stemming from applying th…
We aim to construct the optimal solutions to the undiscounted continuous-time infinite horizon optimization problems, the objective functionals of which may be unbounded. We identify the condition under which the limit of the solutions to the finite horizon problems is optimal for the infinite horizon problems under th…
Optimizes investment under uncertain time horizons with non-concave utility.
A new Bayesian method optimizes time-dependent expensive functions with lookahead.
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
Study of isometries in spacetimes without observer horizons.
We consider the problem of portfolio optimization in a simple incomplete market and under a general utility function. By working with the associated Hamilton-Jacobi-Bellman partial differential equation (HJB PDE), we obtain a closed-form formula for a trading strategy which approximates the optimal trading strategy whe…
Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.
New algorithm learns optimal policies with just 1 episode, settling horizon-dependence in RL.
Study examines Wang-Yau quasi-local energy in strong fields near apparent horizons.
UCRL2-VTR achieves nearly optimal regret for learning MDPs with linear function approximation.
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long run. Yet, it may be difficult (or even intractable) mathematically to learn with th…
The paper finds the shortest time to exploit arbitrage in multi-stock markets.
New research shows exponential lower bounds for planning in MDPs with linearly-realizable optimal action-value functions.
Heterotic horizons preserving 4 supersymmetries have sections which are T^2 fibrations over 6-dimensional conformally balanced Hermitian manifolds. We give new examples of horizons with sections S^3 X S^3 X T^2 and SU(3). We then examine the heterotic horizons which are T^4 fibrations over a Kahler 4-dimensional manifo…
C-Learning estimates reachability over time to solve multi-goal tasks.
The paper optimizes portfolios in a financial market with correlated assets using a stochastic volatility model.
In reinforcement learning, the discount factor controls the agent's effective planning horizon. Traditionally, this parameter was considered part of the MDP; however, as deep reinforcement learning algorithms tend to become unstable when the effective planning horizon is long, recent works refer to as a hyper-p…
Online learning rbfnet improves multi-horizon returns forecasts for financial time series.
Detecting a specific horizon in seismic images is a valuable tool for geological interpretation. Because hand-picking the locations of the horizon is a time-consuming process, automated computational methods were developed starting three decades ago. Older techniques for such picking include interpolation of control po…
In this paper, we propose to combine imitation and reinforcement learning via the idea of reward shaping using an oracle. We study the effectiveness of the near-optimal cost-to-go oracle on the planning horizon and demonstrate that the cost-to-go oracle shortens the learner's planning horizon as function of its accurac…
A new ML algorithm solves complex economic control problems.
Careful tuning of the learning rate, or even schedules thereof, can be crucial to effective neural net training. There has been much recent interest in gradient-based meta-optimization, where one tunes hyperparameters, or even learns an optimizer, in order to minimize the expected loss when the training procedure is un…
In an incomplete market, with incompleteness stemming from stochastic factors imperfectly correlated with the underlying stocks, we derive representations of homothetic (power, exponential and logarithmic) forward performance processes in factor-form using ergodic BSDE. We also develop a connection between the forward …
This study assesses the influence of the forecast horizon on the forecasting performance of several machine learning techniques. We compare the fo recast accuracy of Support Vector Regression (SVR) to Neural Network (NN) models, using a linear model as a benchmark. We focus on international tourism demand to all sevent…
We aim to generalize the results of Cai and Nitta (2007) by allowing both the utility and production function to depend on time. We also consider an additional intertemporal optimality criterion. We clarify the conditions under which the limit of the solutions for the finite horizon problems is optimal among all attain…
New algorithm for RL with horizon-free reward-free exploration for linear MDPs.
We consider generic static spacetimes with Killing horizons and study properties of curvature tensors in the horizon limit. It is determined that the Weyl, Ricci, Riemann and Einstein tensors are algebraically special and mutually aligned on the horizon. It is also pointed out that results obtained in the tetrad adjust…
Using high frequency data, we have studied empirically the change of volatility, also called volatility derivative, for various time horizons. In particular, the correlation between the volatility derivative and the volatility realized in the next time period is a measure of the response function of the market particip…
We present a fully nonparametric method to estimate the value function, via simulation, in the context of expected infinite-horizon discounted rewards for Markov chains. Estimating such value functions plays an important role in approximate dynamic programming and applied probability in general. We incorporate "soft in…
Modeling risk and performance with Levy-stable distributions.
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
MQF forecasts multivariate quantiles globally.
Langevin dynamics fails to produce accurate samples even with small score function errors.
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al. (2018a) proposed an approach that significantly reduces the variance of infinite-horizon off-policy evaluation by estimating the stationar…
Optimizes portfolio in volatile markets with jumps, providing accurate formulas.
New algorithms reduce contextual bandits' regret without knowing reward noise variances.
Scalar-tensor gravitation theories, such as the Brans-Dicke family of theories, are commonly partly described by a modified Einstein equation in which the Ricci tensor is replaced by the Bakry-Émery-Ricci tensor of a Lorentzian metric and scalar field. In physics this formulation is sometimes referred to as the "Jordan…
In this paper, we study optimal liquidation problems in a randomly-terminated horizon. We consider the liquidation of a large single-asset portfolio with the aim of minimizing a combination of volatility risk and transaction costs arising from permanent and temporary market impact. Three different scenarios are analyze…
New algorithms for learning MDPs with linear approximations in infinite-horizon settings.
We analyze (the harmonic map representation of) static solutions of the Einstein Equations in dimension three from the point of view of comparison geometry. We find simple monotonic quantities capturing sharply the influence of the Lapse function on the focussing of geodesics. This allows, in particular, a sharp estima…
Algorithm reduces episode count for CMDPs with constraints.
Study optimal consumption and investment for investors with Epstein-Zin preferences.
Optimal reinsurance and dividend strategy for insurance companies in a finite time.
Proves rigidity of extremal Kerr-Newman horizons.
We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…