The paper targets optimal interventions for long-term outcomes using imputed data and policy learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a model-free algorithm for CMDPs with long-term constraints, achieving optimal regret bounds.
This paper discusses the sensitivity of the long-term expected utility of optimal portfolios for an investor with constant relative risk aversion. Under an incomplete market given by a factor model, we consider the utility maximization problem with long-time horizon. The main purpose is to find the long-term sensitivit…
Improved genetic algorithm optimizes SVR for robust long-term stock index forecasting.
Paper optimizes recommendation systems for long-term business metrics.
For a stochastic factor model we maximize the long-term growth rate of robust expected power utility with parameter . Using duality methods the problem is reformulated as an infinite time horizon, risk-sensitive control problem. Our results characterize the optimal growth rate, an optimal long-term trading s…
Industrial-scale podcast recommender system optimizes long-term listening journeys.
The optimal strategies for a long-term static investor are studied. Given a portfolio of a stock and a bond, we derive the optimal allocation of the capitols to maximize the expected long-term growth rate of a utility function of the wealth. When the bond has constant interest rate, three models for the underlying stoc…
Paper introduces benchmark-neutral pricing for long-term contracts.
Bayesian optimization for long-term outcomes using fast and slow experiments.
The article is devoted to investigating the application of aggregating algorithms to the problem of the long-term forecasting. We examine the classic aggregating algorithms based on the exponential reweighing. For the general Vovk's aggregating algorithm we provide its generalization for the long-term forecasting. For …
New algorithm optimizes long-term user satisfaction in recommendation systems.
New algorithm optimizes for long-term user satisfaction in delayed reward settings.
New algorithm optimizes online network resource allocation with long-term constraints.
TiDE uses MLP for fast, simple long-term time-series forecasting.
This paper balances short-term and long-term rewards in policy learning.
We consider a Bayesian financial market with one bond and one stock where the aim is to maximize the expected power utility from terminal wealth. The solution of this problem is known, however there are some conjectures in the literature about the long-term behavior of the optimal strategy. In this paper we prove now t…
FPG uses fractional calculus for efficient reinforcement learning with long-term memory.
Investigates fund separations and stability for long-term optimal investments.
Finding optimal policies which maximize long term rewards of Markov Decision Processes requires the use of dynamic programming and backward induction to solve the Bellman optimality equation. However, many real-world problems require optimization of an objective that is non-linear in cumulative rewards for which dynami…
We consider online optimization in the 1-lookahead setting, where the objective does not decompose additively over the rounds of the online game. The resulting formulation enables us to deal with non-stationary and/or long-term constraints , which arise, for example, in online display advertising problems. We propose a…
Paper introduces new risk measures for Kelly criterion.
This paper proposes a framework to predict long-term trends and short-term fluctuations in multivariate time series.
We propose a long term portfolio management method which takes into account a liability. Our approach is based on the LQG (Linear, Quadratic cost, Gaussian) control problem framework and then the optimal portfolio strategy hedges the liability by directly tracking a benchmark process which represents the liability. Two…
Most practical recommender systems focus on estimating immediate user engagement without considering the long-term effects of recommendations on user behavior. Reinforcement learning (RL) methods offer the potential to optimize recommendations for long-term user engagement. However, since users are often presented with…
New framework for choosing optimal proxy metrics from past experiments.
This paper studies the long-term growth rate of expected utility from holding a leveraged exchanged-traded fund (LETF), which is a constant proportion portfolio of the reference asset. Working with the power utility function, we develop an analytical approach that employs martingale extraction and involves finding the …
Long term optimal investment problems are studied in a factor model with matrix valued state variables. Explicit parameter restrictions are obtained under which, for an isoelastic investor, the finite horizon value function and optimal strategy converge to their long-run counterparts as the investment horizon approache…
Study examines how risk tolerance impacts long-term investment returns.
This paper analyzes the robust growth rate of leveraged ETFs under uncertain parameters.
This study presents a long-term alternative formula for stock price variation described by a geometric Brownian motion on the basis of median instead of mean or expected values. The proposed method is motivated by the observation made in remote fields, where optimality of bet-hedging or diversification strategies is ex…
New algorithm tackles nonstationary linear bandits with latent dynamics.
In this paper, we assume an insure is allowed to purchase proportional reinsurance and can invest his or her wealth into the financial market where a savings account, stocks and bonds are available. Different from classical optimal investment and reinsurance problem, this paper studies the insurer's long-term investmen…
Stochastic gradient descent's long-term fluctuations are described by a diffusion limit.
We introduce here for the first time the long-term swap rate, characterised as the fair rate of an overnight indexed swap with infinitely many exchanges. Furthermore we analyse the relationship between the long-term swap rate, the long-term yield, see Biagini et al. [2018], Biagini and Härtel [2014], and El Karoui et a…
Kernel method estimates long-term effects from short-term data.
Deep RL drone trained to compete against classical path planning in drone racing.
In this paper, we derive a new model of synaptic plasticity, based on recent algorithms for reinforcement learning (in which an agent attempts to learn appropriate actions to maximize its long-term average reward). We show that these direct reinforcement learning algorithms also give locally optimal performance for the…
We study an exploration method for model-free RL that generalizes the counter-based exploration bonus methods and takes into account long term exploratory value of actions rather than a single step look-ahead. We propose a model-free RL method that modifies Delayed Q-learning and utilizes the long-term exploration bonu…
The paper develops diverse risk models for US stock portfolios.
Open-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals, and training on standard movie or online datasets may lead to the generation of inappropriate, biased, or offensive text. Reinforcement Lear…
Deep RNNs excel at capturing long-term dependencies in sequential data.
Optimizes long-term social welfare in recommender systems by matching users to providers.
Ranking models are typically designed to provide rankings that optimize some measure of immediate utility to the users. As a result, they have been unable to anticipate an increasing number of undesirable long-term consequences of their proposed rankings, from fueling the spread of misinformation and increasing polariz…
Model combines long-term and short-term memory using conceptors.
Investigates long-term performance of multi-fidelity Bayesian optimization.
Efficient algorithm predicts unknown linear systems with long-term memory.
This paper considers online convex optimization over a complicated constraint set, which typically consists of multiple functional constraints and a set constraint. The conventional online projection algorithm (Zinkevich, 2003) can be difficult to implement due to the potentially high computation complexity of the proj…