A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
The policy gradient theorem is defined based on an objective with respect to the initial distribution over states. In the discounted case, this results in policies that are optimal for one distribution over initial states, but may not be uniformly optimal for others, no matter where the agent starts from. Furthermore, …
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…
There is an observed basis between repo discounting, implied from market repo rates, and bond discounting, stripped from the market prices of the underlying bonds. Here, this basis is explained as a convexity effect arising from the decorrelation between the discount rates for derivatives and bonds. Using a Hull-White …
New RL approach handles non-exponential discounting for sequential decisions.
problem Modeling human discounting in sequential decision-making tasks.
method Generalized model-based reinforcement learning with arbitrary discount functions, using Hamilton-Jacobi-Bellman equation and collocation method.
result Validated approach on simulated problems, showing applicability to human discounting.
The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …
We show that different rates should be used for borrowing and discount rates, and that the risk-free rate should be used for discounting when assessing and comparing the cost of energy accross diffferent producers and technologies, on the example of photovoltaics. Recent quantitative models using the same rate for borr…
We consider a discounted reward control problem in continuous time stochastic environment where the discount rate might be an unbounded function of the control process. We provide a set of general assumptions to ensure that there exists a smooth classical solution to the corresponding HJB equation. Moreover, some verif…
Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…
In this paper, we study the dividend strategies for a shareholder with non-constant discount rate in a diffusion risk model. We assume that the dividends can only be paid at a bounded rate and restrict ourselves to the Markov strategies. This is a time inconsistent control problem. The extended HJB equation is given an…
In this paper we extend the existing literature on xVA along three directions. First, we enhance current BSDE-based xVA frameworks to include initial margin in presence of defaults. Next, we solve the consistency problem that arises when the front-office desk of the bank uses trade-specific discount curves (CSA discoun…
This paper considers the problem of consumption and investment in a financial market within a continuous time stochastic economy. The investor exhibits a change in the discount rate. The investment opportunities are a stock and a riskless account. The market coefficients and discount factor switch according to a finite…
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the discount factor from the state distribution and therefore do not optimize the discounted objective. What do they optimize instead? This has be…
We prove that a large class of discrete-time insurance surplus processes converge weakly to a generalized Ornstein-Uhlenbeck process, under a suitable re-normalization and when the time-step goes to 0. Motivated by ruin theory, we use this result to obtain approximations for the moments, the ultimate ruin probability a…
We consider a generalization of the Heath Jarrow Morton model for the term structure of interest rates where the forward rate is driven by Paretian fluctuations. We derive a generalization of Itô's lemma for the calculation of a differential of a Paretian stochastic variable and use it to derive a Stochastic Differenti…
Intertemporal decision making involves choices among options whose effects occur at different moments. These choices are influenced not only by the effect of rewards value perception at different moments, but also by the time perception effect. One of the main difficulties that affect standard experiments involving int…