New method uses entropy to improve policy gradient exploration.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New RL approach uses future state and action visitation measures for better exploration.
Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
Proves lower discount rates are needed for future losses.
Model shows how discount rates affect intergenerational equity in climate mitigation.
The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …
New findings reveal discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
This paper improves Thompson Sampling for complex decision-making problems.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
We introduce a simple generalization of rational bubble models which removes the fundamental problem discovered by [Lux and Sornette, 1999] that the distribution of returns is a power law with exponent less than 1, in contradiction with empirical data. The idea is that the price fluctuations associated with bubbles mus…
For environmental problems such as global warming future costs must be balanced against present costs. This is traditionally done using an exponential function with a constant discount rate, which reduces the present value of future costs. The result is highly sensitive to the choice of discount rate and has generated …
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…
In this paper, we settle the sampling complexity of solving discounted two-player turn-based zero-sum stochastic games up to polylogarithmic factors. Given a stochastic game with discount factor we provide an algorithm that computes an -optimal strategy with high-probability given $\tilde{O}((1 - γ)^{-3}…
We propose an analytically tractable variation of the minority game in which rational agents use probabilistic strategies. In our model, agents choose between two alternatives repeatedly, and those who are in the minority get a pay-off 1, others zero. The agents optimize the expectation value of their discounted fu…
Stochastic dividend discount models (Hurley and Johnson, 1994 and 1998, Yao, 1997) present expressions for the expected value of stock prices when future dividends evolve according to some random scheme. In this paper we try to offer a more precise view on this issue proposing a closed-form formula for the variance of …
Paper introduces non-linear discounting models for default compensation and climate valuation.
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters. However, most policy gradient methods drop the discount factor from the state distribution and therefore do not optimize the discounted objective. What do they optimize instead? This has be…
RegFlow models future states with flexible probability distributions.
In this paper we study the valuation problem of an insurance company by maximizing the expected discounted future dividend payments in a model with partial information that allows for a changing economic environment. The surplus process is modeled as a Brownian motion with drift. This drift depends on an underlying Mar…
There is a consensus that human and non-human subjects experience temporal distortions in many stages of their perceptual and decision-making systems. Similarly, intertemporal choice research has shown that decision-makers undervalue future outcomes relative to immediate ones. Here we combine techniques from informatio…
The study constructs models for SOFR term rates using futures data.
We investigate the problem of optimal dividend distribution for a company in the presence of regime shifts. We consider a company whose cumulative net revenues evolve as a Brownian motion with positive drift that is modulated by a finite state Markov chain, and model the discount rate as a deterministic function of the…
In many real-world reinforcement learning applications, access to the environment is limited to a fixed dataset, instead of direct (online) interaction with the environment. When using this data for either evaluation or training of a new policy, accurate estimates of discounted stationary distribution ratios -- correct…
Study reveals a hidden cost in derivatives markets through option-implied discount factors.
A firm with heterogeneous shareholders optimizes dividends under ambiguity aggregation.
New RL difficulty shown for discounted settings.
This work improves policy optimization by maximizing entropy of state distribution, leading to better exploration.
Empirical study on long-term discount rates using historical bond prices.
DS-TS adapts to abrupt and smooth changes in bandit problems.
High future discounting rates favor inaction on present expending while lower rates advise for a more immediate political action. A possible approach to this key issue in global economy is to take historical time series for nominal interest rates and inflation, and to construct then real interest rates and finally obta…
The paper reviews historical and modern approaches to asset pricing probability measures.
We consider the valuation problem of an (insurance) company under partial information. Therefore we use the concept of maximizing discounted future dividend payments. The firm value process is described by a diffusion model with constant and observable volatility and constant but unknown drift parameter. For transformi…
The paper models stochastic interest rates for life insurance using phase-type distributions.
In reinforcement learning, Return, which is the weighted accumulated future rewards, and Value, which is the expected return, serve as the objective that guides the learning of the policy. In classic RL, return is defined as the exponentially discounted sum of future rewards. One key insight is that there could be many…
Modeling climate change costs with stochastic interest rates shows inequality, but funding abatement can reduce this.
New RL method improves on standard discounted RL for operations research.
Study optimal stopping times for multi-dimensional processes with non-exponential discounting.
A major challenge in reinforcement learning (RL) is the design of agents that are able to generalize across tasks that share common dynamics. A viable solution is meta-reinforcement learning, which identifies common structures among past tasks to be then generalized to new tasks (meta-test). In meta-training, the RL ag…
Siegel's paradox is a fundamental question in international finance about exchange rates for futures contracts and has puzzled many scholars for over forty years. The unorthodox approach presented in this article leads to an arbitrage-free solution which is invariant under currency re-denominations and is symmetric, as…
This paper considers the problem of consumption and investment in a financial market within a continuous time stochastic economy. The investor exhibits a change in the discount rate. The investment opportunities are a stock and a riskless account. The market coefficients and discount factor switch according to a finite…
UCBVI-γ algorithm minimizes regret in discounted MDPs.
NVMDP framework tackles non-stationary MDPs with varying discount rates.
New RL approach handles non-exponential discounting for sequential decisions.
A pricing formula for discount bonds, based on the consideration of the market perception of future liquidity risk, is established. An information-based model for liquidity is then introduced, which is used to obtain an expression for the bond price. Analysis of the bond price dynamics shows that the bond volatility is…
In the "positive interest" models of Flesaker-Hughston, the nominal discount bond system is determined by a one-parameter family of positive martingales. In the present paper we extend this analysis to include a variety of distributions for the martingale family, parameterised by a function that determines the behaviou…
Adaptive online learning algorithm improves history forgetting in nonstationary environments.
Lower discount factors act as a regularizer in RL, improving performance.