Lower discount factors act as a regularizer in RL, improving performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New findings reveal discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …
Entropy-regularized NPG methods converge linearly in discounted MDPs.
New method for optimistic planning in MDPs using regularization.
Adaptive online learning algorithm improves history forgetting in nonstationary environments.
Framework for transferring discount curve estimates across fixed-income product classes.
Deep neural networks decompose SDF into linear and nonlinear components.
Paper develops a discounted algorithm for online convex optimization that adapts to unknown discount factors.
Paper introduces non-linear discounting models for default compensation and climate valuation.
We propose and study a general framework for regularized Markov decision processes (MDPs) where the goal is to find an optimal policy that maximizes the expected discounted total reward plus a policy regularization term. The extant entropy-regularized MDPs can be cast into our framework. Moreover, under our framework, …
Study analyzes how discounts affect train ticket purchases and rescheduling in Switzerland.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
Proves lower discount rates are needed for future losses.
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…
There is an observed basis between repo discounting, implied from market repo rates, and bond discounting, stripped from the market prices of the underlying bonds. Here, this basis is explained as a convexity effect arising from the decorrelation between the discount rates for derivatives and bonds. Using a Hull-White …
Proposes a new framework for discount models.
For an infinite-horizon continuous-time optimal stopping problem under non-exponential discounting, we look for an optimal equilibrium, which generates larger values than any other equilibrium does on the entire state space. When the discount function is log sub-additive and the state process is one-dimensional, an opt…
This paper shows how forward rate interpolations are equivalent to discount factor interpolations in yield curve construction.
Entropy regularization is used to get improved optimization performance in reinforcement learning tasks. A common form of regularization is to maximize policy entropy to avoid premature convergence and lead to more stochastic policies for exploration through action space. However, this does not ensure exploration in th…
New RL approach handles non-exponential discounting for sequential decisions.
The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …
We optimize discounts to maximize influence spread in social networks.
Study optimal portfolio strategies with time-varying discount rates.
The study uses reproducing kernels to model bond discount curves.
We show that different rates should be used for borrowing and discount rates, and that the risk-free rate should be used for discounting when assessing and comparing the cost of energy accross diffferent producers and technologies, on the example of photovoltaics. Recent quantitative models using the same rate for borr…
We consider a discounted reward control problem in continuous time stochastic environment where the discount rate might be an unbounded function of the control process. We provide a set of general assumptions to ensure that there exists a smooth classical solution to the corresponding HJB equation. Moreover, some verif…
New RL difficulty shown for discounted settings.
UCBVI-γ algorithm minimizes regret in discounted MDPs.
Derivative pricing is about cash flow discounting at the riskfree rate. This teaching has lost its meaning post the financial crisis, due to the addition of extra value adjustments (XVA), which also made derivatives pricing and valuation a very difficult task for investors. This article recovers a properly defined disc…
Study optimal stopping problems with finite-time horizon and proves continuity and strict monotonicity of the boundary.
Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…
In this paper, we study the dividend strategies for a shareholder with non-constant discount rate in a diffusion risk model. We assume that the dividends can only be paid at a bounded rate and restrict ourselves to the Markov strategies. This is a time inconsistent control problem. The extended HJB equation is given an…
Paper proposes a machine learning method to predict sale efficacy.
In this paper we extend the existing literature on xVA along three directions. First, we enhance current BSDE-based xVA frameworks to include initial margin in presence of defaults. Next, we solve the consistency problem that arises when the front-office desk of the bank uses trade-specific discount curves (CSA discoun…
Study improves dividend discount model using VAR process.
Study optimal stopping for group with diverse discount rates using an attitude function.
Study optimal stopping times for multi-dimensional processes with non-exponential discounting.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
Proposes a new method for determining LGD discount rates based on cost of capital.
Paper tackles time inconsistency in portfolio management with stochastic volatility and power utility.
This paper considers the problem of consumption and investment in a financial market within a continuous time stochastic economy. The investor exhibits a change in the discount rate. The investment opportunities are a stock and a riskless account. The market coefficients and discount factor switch according to a finite…
Optimal online linear regression in dynamic environments using discounted Vovk-Azoury-Warmuth forecaster.
Softmax policy gradient achieves global optimality in wide neural networks with entropy regularization.
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the discount factor from the state distribution and therefore do not optimize the discounted objective. What do they optimize instead? This has be…
Study on mean field games with singular controls and their applications.
GPMD solves regularized RL with linear convergence, promoting structural policies.
Intertemporal decision making involves choices among options whose effects occur at different moments. These choices are influenced not only by the effect of rewards value perception at different moments, but also by the time perception effect. One of the main difficulties that affect standard experiments involving int…