Improved sample complexity for identifying best policies in risk-sensitive reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and are optimal if and only if the latter is exchangeable, and if and only if the opti…
New research shows exponential lower bounds for planning in MDPs with linearly-realizable optimal action-value functions.
Optimal execution of portfolio transactions is the essential part of algorithmic trading. In this paper we present in simple analytical form the optimal trajectory for risk-averse trader with the assumption of exponential market recovery and short-time investment horizon.
In this paper, we obtain analytical expression for the distribution of the occupation time in the red (below level ) up to an (independent) exponential horizon for spectrally negative Lévy risk processes and refracted spectrally negative Lévy risk processes. This result improves the existing literature in which only…
We present a continuous-time maximum likelihood estimation methodology for credit rating transition probabilities, taking into account the presence of censored data. We perform rolling estimates of the transition matrices with exponential time weighting with varying horizons and discuss the underlying dynamics of trans…
In an incomplete market, with incompleteness stemming from stochastic factors imperfectly correlated with the underlying stocks, we derive representations of homothetic (power, exponential and logarithmic) forward performance processes in factor-form using ergodic BSDE. We also develop a connection between the forward …
We study the probability distribution of stock returns at mesoscopic time lags (return horizons) ranging from about an hour to about a month. While at shorter microscopic time lags the distribution has power-law tails, for mesoscopic times the bulk of the distribution (more than 99% of the probability) follows an expon…
We study the optimal stopping of an American call option in a random time-horizon under exponential spectrally negative Lévy models. The random time-horizon is modeled as the so-called Omega default clock in insurance, which is the first time when the occupation time of the underlying Lévy process below a level , ex…
Consider power utility maximization of terminal wealth in a 1-dimensional continuous-time exponential Levy model with finite time horizon. We discretize the model by restricting portfolio adjustments to an equidistant discrete time grid. Under minimal assumptions we prove convergence of the optimal discrete-time strate…
SMRL uses score matching for efficient RL with exponential family models.
Mirror flow optimizes separable data problems, converging to a maximum margin classifier.
In this paper, we investigate the Merton portfolio management problem in the context of non-exponential discounting. This gives rise to time-inconsistency of the decision-maker. If the decision-maker at time t=0 can commit his/her successors, he/she can choose the policy that is optimal from his/her point of view, and …
Unique optimal strategy identified for state-dependent risk aversion.
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes. We analyze the variance of the most popular approaches through the viewpoint of con…
New insights into how to inspect and learn from multi-stage processes and AI reasoning.
Behavior cloning training instabilities amplified by SGD noise over long horizons.
An online reinforcement learning algorithm is anytime if it does not need to know in advance the horizon T of the experiment. A well-known technique to obtain an anytime algorithm from any non-anytime algorithm is the "Doubling Trick". In the context of adversarial or stochastic multi-armed bandits, the performance of …
TensorPlan algorithm finds δ-optimal policies with poly queries under linearly realizable state-value function.
New algorithm reduces reinforcement learning complexity, approaching contextual bandits.
Consider a discrete-time infinite horizon financial market model in which the logarithm of the stock price is a time discretization of a stochastic differential equation. Under conditions different from those given in a previous paper of ours, we prove the existence of investment opportunities producing an exponentiall…
TensorPlan shows an exponential lower bound for planning in MDPs with linearly realizable value functions.
FMDP-BF algorithm improves RL in factored MDPs with exponential regret reduction.
Long term optimal investment problems are studied in a factor model with matrix valued state variables. Explicit parameter restrictions are obtained under which, for an isoelastic investor, the finite horizon value function and optimal strategy converge to their long-run counterparts as the investment horizon approache…
In this paper a quantitative analysis of the ruin probability in finite time of discrete risk process with proportional reinsurance and investment of finance surplus is focused on. It is assumed that the total loss on a unit interval has a light-tailed distribution -- exponential distribution and a heavy-tailed distrib…
New Thompson sampling algorithm reduces regret for exponential family bandits.
Logarithmic regret achieved in continuous-time linear-quadratic reinforcement learning.
Softmax PG methods can take extremely long to converge, even with exact gradients.
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…
Logarithmic regret for continuous-time reinforcement learning.
We extend the model of stochastic bandits with adversarial corruption (Lykouriset al., 2018) to the stochastic linear optimization problem (Dani et al., 2008). Our algorithm is agnostic to the amount of corruption chosen by the adaptive adversary. The regret of the algorithm only increases linearly in the amount of cor…
Paper provides exponential convergence guarantees for Iterative Markovian Fitting.
New algorithm for reward-free RL with linear function approximation, reducing sample complexity.
We give a new proof of the fact that the value function of the finite time horizon American put option for a jump diffusion, when the jumps are from a compound Poisson process, is the classical solution of a free boundary equation. We also show that the value function is across the optimal stopping boundary. Our …
We construct a large class of dynamical vacuum black hole spacetimes whose exterior geometry asymptotically settles down to a fixed Schwarzschild or Kerr metric. The construction proceeds by solving a backwards scattering problem for the Einstein vacuum equations with characteristic data prescribed on the event horizon…
Decentralized algorithm reduces regret and converges to Nash equilibrium in online congestion games.
Study on massless Vlasov equation on Reissner-Nordström spacetimes, showing decay rates and non-decay phenomena.
The paper improves model-based reinforcement learning by using multi-timestep objectives.
Continuous control imitation learning fails if expert actions are smooth.
We study the problem of controlling linear time-invariant systems with known noisy dynamics and adversarially chosen quadratic losses. We present the first efficient online learning algorithms in this setting that guarantee regret under mild assumptions, where is the time horizon. Our algorithms rely …
We present an approach for pricing European call options in presence of proportional transaction costs, when the stock price follows a general exponential Lévy process. The model is a generalization of the celebrated work of Davis, Panas and Zariphopoulou (1993), where the value of the option is defined as the utility …
This paper derives a portfolio decomposition formula when the agent maximizes utility of her wealth at some finite planning horizon. The financial market is complete and consists of multiple risky assets (stocks) plus a risk free asset. The stocks are modelled as exponential Brownian motions with drift and volatility b…
New algorithms learn MDPs with better regret bounds using generative sampling.
New algorithm reduces regret in private online learning with optimal gap-dependent rate.
Optimizes liquidation strategies for assets with Levy process price dynamics.
Study uses G-BSDEs to decompose pricing kernels under robust G-expectation.
This paper analyzes a simplified strategy for nonlinear control using local linear models and iLQR updates.
Quantum UCB algorithm reduces reinforcement learning regret exponentially.