Improved off-policy reinforcement learning by discounting and soft normalization.
problem Divergence issues in off-policy reinforcement learning.
method Introducing a discount factor and a soft normalization penalty into COP-TD.
result Discounted COP-TD is better behaved both theoretically and empirically.
The problem of on-line off-policy evaluation (OPE) has been actively studied in the last decade due to its importance both as a stand-alone problem and as a module in a policy improvement scheme. However, most Temporal Difference (TD) based solutions ignore the discrepancy between the stationary distribution of the beh…
Paper develops a discounted algorithm for online convex optimization that adapts to unknown discount factors.
problem Developing an algorithm that can adapt to an unknown discount factor in online convex optimization.
method Smoothed Online Gradient Descent (SOGD) with Discounted-Normal-Predictor (DNP).
result Achieves a uniform O ( log T / 1 − λ ) O(\sqrt{\log T/1-λ}) O ( log T /1 − λ ) discounted regret across a continuous interval of discount factors. Repo rates are explained as a convexity effect from bond and derivative discount rates.
problem Explaining the observed basis between repo rates and bond prices.
method Using a Hull-White model, derived expressions for repo rates and extrapolation.
result Interpolated and extrapolated repo curves for bond-collateralised derivatives.
Paper introduces non-linear discounting models for default compensation and climate valuation.
problem Valuation of non-replicable value and damage under default risk.
method Develops two models: one for risk-neutralising discounting and another for survival probability dependent discounting.
result Non-decaying discount factors (negative discount rates) are possible under certain scenarios.
This work bridges hyperbolic discounting in RL with exponential discounting.
problem Hyperbolic discounting in reinforcement learning models.
method Implemented a hyperbolic discounting RL agent and demonstrated its effectiveness.
result Hyperbolic discounting can be approximated using familiar RL techniques.
Study analyzes how discounts affect train ticket purchases and rescheduling in Switzerland.
problem Understanding how discounts influence train ticket buying and rescheduling behavior.
method Machine learning techniques, including causal machine learning, to analyze survey data.
result Increasing a discount rate by 1% increases the rescheduled trip share by 0.16% among always buyers.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
Proves lower discount rates are needed for future losses.
problem Determining appropriate discount rates for future losses.
method Analyzes climate change and discount rates debate.
result Risk requires a lower, not higher, discount rate.
Proposes a new framework for discount models.
problem Arbitrage-free dynamic framework for discount models.
method Derives general consistency conditions for factor models.
result Alternative to Heath--Jarrow--Morton framework for forward rates.
Revises derivative pricing post financial crisis by defining a discount rate.
problem Derivative pricing became complex with XVA adjustments.
method Developed a binomial tree model for pricing with counterparty and funding risks.
result Coherent XVAs naturally result from decomposing the discount rate.
Enhances xVA valuation with initial margin and CSA discounting.
problem Inconsistency in discount curves and aggregation of trades.
method Unified approach using BSDEs with initial margin, CSA discounting, and non-linear aggregation.
result Consistent xVA valuation with trade-specific discount curves.
This paper shows how forward rate interpolations are equivalent to discount factor interpolations in yield curve construction.
problem The challenge of choosing between different interpolation methods for yield curve construction.
method Demonstrates the equivalence between forward rate interpolations and discount factor interpolations.
result Some popular interpolation methods on forward rates are equivalent to classical interpolation methods on discount factors.
New RL approach handles non-exponential discounting for sequential decisions.
problem Modeling human discounting in sequential decision-making tasks.
method Generalized model-based reinforcement learning with arbitrary discount functions, using Hamilton-Jacobi-Bellman equation and collocation method.
result Validated approach on simulated problems, showing applicability to human discounting.
The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …
We optimize discounts to maximize influence spread in social networks.
problem Maximizing influence spread in social networks with fractional discounts.
method Developed an efficient (1-1/e)-approximation algorithm for NP-hard problem.
result Achieved an approximation of 1-1/e for influence maximization.
Study optimal portfolio strategies with time-varying discount rates.
problem Optimizing portfolio decisions with a non-constant discount rate.
method Introduced subgame perfect strategies to handle time inconsistency, using fixed point iteration to find the utility-weighted discount rate.
result Subgame perfect strategies are equivalent to optimal strategies under certain utility function assumptions.
The study uses reproducing kernels to model bond discount curves.
problem Estimating bond discount curves under no-arbitrage conditions.
method Introduced reproducing kernels as a regression basis for estimating bond discount curves.
result Reproducing kernels provide a tractable solution for calibrating models to market data.
Policy gradient methods do not optimize the discounted objective, leading to suboptimal results.
problem Understanding the true optimization objective of policy gradient methods.
method Analyzing the update direction of policy gradient methods and proving it is not the gradient of any function.
result Policy gradient methods do not optimize the discounted objective, leading to suboptimal results.
We show that different rates should be used for borrowing and discount rates, and that the risk-free rate should be used for discounting when assessing and comparing the cost of energy accross diffferent producers and technologies, on the example of photovoltaics. Recent quantitative models using the same rate for borr…
New method uses entropy to improve policy gradient exploration.
problem Limited exploration in policy gradient methods.
method Entropy regularization with discounted future state distribution.
result Proves convergence to locally optimal policy.
We consider a discounted reward control problem in continuous time stochastic environment where the discount rate might be an unbounded function of the control process. We provide a set of general assumptions to ensure that there exists a smooth classical solution to the corresponding HJB equation. Moreover, some verif…
The paper proposes a method to discount backtest PnLs due to in-sample overfitting.
problem In-sample overfitting in backtest-based investment strategies.
method A simple framework to model and quantify in-sample PnL overfitting.
result Computes the appropriate discount factor for PnLs of in-sample investment strategies.
New findings reveal discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
problem Discount regularization leads to poor performance in unevenly sampled data.
method Equivalence theorem showing discount regularization as a strong prior, setting regularization parameters locally for individual state-action pairs.
result Discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
UCBVI-γ algorithm minimizes regret in discounted MDPs.
problem Minimizing regret in discounted MDPs.
method Optimism in the face of uncertainty principle and Bernstein-type bonus.
result UCBVI-γ achieves nearly minimax optimal regret.
New RL difficulty shown for discounted settings.
problem Difficulty in reinforcement learning with discounted rewards.
method Adapted Wang et al. (2020) construction to 2-state MDP.
result Learning impossible even with infinite data in discounted setting.
Lower discount factors act as a regularizer in RL, improving performance.
problem Improving RL performance with limited data.
method Explicitly equating reduced discount factors to regularization terms.
result Regularization effectiveness depends on data properties.
Paper proposes a machine learning method to predict sale efficacy.
problem Determining the efficacy of online sales from discounts alone.
method Machine learning-based heuristic using Support Vector Machine.
result Predicts sale efficacy with 91.11% accuracy.
Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…
In this paper, we study the dividend strategies for a shareholder with non-constant discount rate in a diffusion risk model. We assume that the dividends can only be paid at a bounded rate and restrict ourselves to the Markov strategies. This is a time inconsistent control problem. The extended HJB equation is given an…
A new method maps value estimates to logarithmic space to enable lower discount factors in reinforcement learning.
problem The poor performance of low discount factors in reinforcement learning.
method Introducing a logarithmic mapping to value estimates.
result The method enables lower discount factors, solving challenging reinforcement learning problems.
Study improves dividend discount model using VAR process.
problem Improving dividend discount models for better predictions.
method Introduced a Gordon growth model based on Vector Autoregressive Process (VAR).
result Two Propositions related to the new model.
Study optimal stopping for group with diverse discount rates using an attitude function.
problem Optimal stopping for a group with diverse discount rates under an aggregation preference.
method Develop iterative approach using consistent planning for time-consistent equilibria.
result Characterize all time-consistent mild equilibria as fixed points of an operator.
Study time-inconsistent consumption-investment in incomplete markets with general discount functions.
problem Time-inconsistent consumption-investment problems in incomplete markets.
method Coupled forward-backward stochastic differential equation approach.
result Uniqueness of open-loop equilibrium pair proved.
Study optimal stopping times for multi-dimensional processes with non-exponential discounting.
problem Optimal stopping in multi-dimensional processes with non-exponential discounting.
method Probabilistic potential theory to establish existence of optimal equilibria.
result Existence of optimal equilibria for multi-dimensional stopping problems.
Proposes a new method for determining LGD discount rates based on cost of capital.
problem Determining an appropriate discount rate for LGD estimation.
method Market-consistent pricing of defaulted loan portfolios to infer discount rates.
result Discount rates reflect both undiversifiable risk and time value of money.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
problem Optimal dividend and capital injection strategies under time-inconsistent preferences.
method Diffusion risk model with general discount functions, weak equilibrium definition, HJB equation system.
result Explicit solutions and threshold types of optimal strategies derived under different discount functions.
Paper tackles time inconsistency in portfolio management with stochastic volatility and power utility.
problem Time inconsistency in portfolio management with stochastic volatility and power utility.
method Extended Hamilton Jacobi Bellman (HJB) equation, fixed point iteration, and linear parabolic PDE.
result Subgame perfect strategies are characterized and solved through numerical experiments.
This paper considers the problem of consumption and investment in a financial market within a continuous time stochastic economy. The investor exhibits a change in the discount rate. The investment opportunities are a stock and a riskless account. The market coefficients and discount factor switch according to a finite…
Optimal online linear regression in dynamic environments using discounted Vovk-Azoury-Warmuth forecaster.
problem Achieving optimal performance in dynamic online linear regression without prior knowledge.
method Developed a discounted variant of the Vovk-Azoury-Warmuth forecaster to achieve optimal dynamic regret guarantees.
result Achieved dynamic regret of the form $O\left(d\log(T)\vee \sqrt{dP_{T}^γ(\vec{u})T}
ight)$ , with a learnable discount factor.
Study on mean field games with singular controls and their applications.
problem Optimal productivity expansion in dynamic oligopolies.
method Existence and uniqueness of mean field equilibria through nonlinear equations, Abelian limit for discounted and ergodic games.
result Valid connection between discounted and ergodic games, approximation of Nash equilibria.
Intertemporal decision making involves choices among options whose effects occur at different moments. These choices are influenced not only by the effect of rewards value perception at different moments, but also by the time perception effect. One of the main difficulties that affect standard experiments involving int…
Study on CEF discount in Bangladesh, finds size and maturity impact, turnover negative.
problem Exploring the discount puzzle in closed-end mutual funds in Bangladesh.
method Fixed effects panel regression with diagnostic tests.
result Fund size and maturity positively impact CEF discount, turnover negatively impacts.
Investment decisions shift earlier as patience decreases, with implications for pasting conditions.
problem Investment timing under decreasing impatience.
method Game-theoretic framework with continuous-time capacity expansion problem.
result Decreasing impatience leads to earlier investment decisions, but can violate smooth pasting conditions.
Approximates discounted moments for financial products using polynomial expansions.
problem Approximating discounted moments of stochastic processes for financial applications.
method High-order power series expansion of the infinitesimal generator.
result Error decreases to around 10 to 100 times machine precision for higher orders.
New algorithm reduces online regression error in RKHS.
problem Online regression with time-varying functions in RKHS.
method Hierarchical Vovk-Azoury-Warmuth with discounting.
result Achieves optimal dynamic regret with O ( T 2 / 3 P T 1 / 3 + T ln T ) O(T^{2/3}P_T^{1/3} + \sqrt{T}\ln T) O ( T 2/3 P T 1/3 + T ln T ) regret bound. Empirical study on long-term discount rates using historical bond prices.
problem Estimating long-term real interest rates and discount rates from historical bond data.
method Using Fourier transforms to derive the discount function and fitting it to historical data.
result Estimated long-term discount rates of 1.7% for UK and 2.2% for US.
New Q Q Q -learning method reduces variance and achieves optimal sample complexity.
problem Improving Q Q Q -learning to reduce variance and improve sample efficiency. method Introduces variance-reduced Q Q Q -learning and analyzes its sample complexity. result Achieves minimax optimal sample complexity for estimating optimal Q Q Q -function.