Q-Learning overestimation bias influenced by learning rate, discount factor, and reward signal.
problem Overestimation bias in Q-Learning algorithm.
method Investigated the influence of learning rate, discount factor, and reward signal on Q-Learning's overestimation bias. Tuned parameters and used an exponential moving average of reward signal.
result Q-Learning can achieve more accurate value estimates by tuning parameters and using an exponential moving average of reward signal.
Study on investment strategy for agents with periodic preferences and discounting.
problem Investment decisions by agents with periodic S-shaped preferences and present bias.
method Infinite-horizon, continuous-time portfolio selection problem with quasi-hyperbolic discounting.
result Time-consistent planning strategy can be formulated as an equilibrium to a static mean field game.
This paper solves Siegel's paradox about future exchange rates.
problem Understanding future exchange rates and resolving Siegel's paradox.
method An unorthodox approach leading to an arbitrage-free solution.
result A formula describing all no-arbitrage forward exchange rates.
This paper closely examines theoretical and practical aspects of the widely used discounted cash flows (DCF) valuation method. It assesses its potentials as well as several weaknesses. A special emphasize is being put on the valuation of companies using the DCF method. The paper finds that the discounted cash flow meth…
New objective reduces bias and variance in reinforcement learning derivatives.
problem Estimating derivatives in reinforcement learning with unknown dynamics.
method Derives an objective function compatible with any advantage estimators, allowing trade-off between bias and variance.
result Demonstrates effectiveness in both theoretical and practical settings.
There exist a number of reinforcement learning algorithms which learnby climbing the gradient of expected reward. Their long-runconvergence has been proved, even in partially observableenvironments with non-deterministic actions, and without the need fora system model. However, the variance of the gradient estimator ha…
A new method separates long-term value functions into components for better reinforcement learning.
problem Learning long-term goals in reinforcement learning settings with temporal discounting.
method TD( Δ Δ Δ ) learning, which breaks down value functions into components based on discount factors. result TD( Δ Δ Δ ) learning improves scalability and performance over standard TD learning in certain settings. The paper analyzes the tradeoff between bias and overfitting in reinforcement learning with partial observability.
problem Analyzing the tradeoff between asymptotic bias and overfitting in reinforcement learning with partial observability.
method Theoretical analysis and empirical illustration using truncated history of observations and function approximators.
result A smaller state representation decreases the risk of overfitting, but potentially increases asymptotic bias.
Q( Δ Δ Δ )-Learning improves Q-Learning by separating action-value functions into different time scales.
problem Q-Learning struggles with bias-variance trade-off, especially in long-term rewards.
method Introduces Q( Δ Δ Δ )-Learning, extending TD( Δ Δ Δ ) to decompose Q( Δ Δ Δ )-function into distinct discount factors. result Q( Δ Δ Δ )-Learning achieves better stability and scalability, especially for long-term tasks. We establish explicit socially optimal rules for an irreversible investment deci- sion with time-to-build and uncertainty. Assuming a price sensitive demand function with a random intercept, we provide comparative statics and economic interpreta- tions for three models of demand (arithmetic Brownian, geometric Brownian…
RMC uses renewal theory for online reinforcement learning with low variance and easy implementation.
problem Online reinforcement learning for infinite horizon Markov decision processes.
method RMC combines Monte Carlo methods with renewal theory to estimate performance gradients and update policies.
result RMC converges to locally optimal policies and can be generalized to post-decision state models.
Optimizes learning policies in average-reward MDPs with improved sample complexity.
problem Learning optimal policies in average-reward MDPs with limited samples.
method Reduces to discounted MDPs and uses improved bounds for variance parameters.
result Establishes minimax optimal sample complexity bound of O(SA(H/ε^2))
Optimizes learning policies in MDPs with weakly communicating structure.
problem Learning optimal policies in weakly communicating MDPs with generative model.
method Span-based approach, reducing to discounted MDPs for analysis.
result First minimax optimal sample complexity bound for weakly communicating MDPs.
Paper develops a discounted algorithm for online convex optimization that adapts to unknown discount factors.
problem Developing an algorithm that can adapt to an unknown discount factor in online convex optimization.
method Smoothed Online Gradient Descent (SOGD) with Discounted-Normal-Predictor (DNP).
result Achieves a uniform O ( log T / 1 − λ ) O(\sqrt{\log T/1-λ}) O ( log T /1 − λ ) discounted regret across a continuous interval of discount factors. This work overcomes bias in concave multi-objective reinforcement learning.
problem Gradient bias in policy gradient methods for concave scalarized multi-objective reinforcement learning.
method Developed a Natural Policy Gradient (NPG) algorithm with a multi-level Monte Carlo (MLMC) estimator.
result Achieved optimal O ~ ( ε − 2 ) \widetilde{\mathcal{O}}(ε^{-2}) O ( ε − 2 ) sample complexity for computing an ε ε ε -optimal policy. Repo rates are explained as a convexity effect from bond and derivative discount rates.
problem Explaining the observed basis between repo rates and bond prices.
method Using a Hull-White model, derived expressions for repo rates and extrapolation.
result Interpolated and extrapolated repo curves for bond-collateralised derivatives.
Paper introduces non-linear discounting models for default compensation and climate valuation.
problem Valuation of non-replicable value and damage under default risk.
method Develops two models: one for risk-neutralising discounting and another for survival probability dependent discounting.
result Non-decaying discount factors (negative discount rates) are possible under certain scenarios.
This work bridges hyperbolic discounting in RL with exponential discounting.
problem Hyperbolic discounting in reinforcement learning models.
method Implemented a hyperbolic discounting RL agent and demonstrated its effectiveness.
result Hyperbolic discounting can be approximated using familiar RL techniques.
Study analyzes how discounts affect train ticket purchases and rescheduling in Switzerland.
problem Understanding how discounts influence train ticket buying and rescheduling behavior.
method Machine learning techniques, including causal machine learning, to analyze survey data.
result Increasing a discount rate by 1% increases the rescheduled trip share by 0.16% among always buyers.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
Proves lower discount rates are needed for future losses.
problem Determining appropriate discount rates for future losses.
method Analyzes climate change and discount rates debate.
result Risk requires a lower, not higher, discount rate.
Proposes a new framework for discount models.
problem Arbitrage-free dynamic framework for discount models.
method Derives general consistency conditions for factor models.
result Alternative to Heath--Jarrow--Morton framework for forward rates.
Revises derivative pricing post financial crisis by defining a discount rate.
problem Derivative pricing became complex with XVA adjustments.
method Developed a binomial tree model for pricing with counterparty and funding risks.
result Coherent XVAs naturally result from decomposing the discount rate.
Enhances xVA valuation with initial margin and CSA discounting.
problem Inconsistency in discount curves and aggregation of trades.
method Unified approach using BSDEs with initial margin, CSA discounting, and non-linear aggregation.
result Consistent xVA valuation with trade-specific discount curves.
This paper shows how forward rate interpolations are equivalent to discount factor interpolations in yield curve construction.
problem The challenge of choosing between different interpolation methods for yield curve construction.
method Demonstrates the equivalence between forward rate interpolations and discount factor interpolations.
result Some popular interpolation methods on forward rates are equivalent to classical interpolation methods on discount factors.
Optimizes spending by adjusting a discount factor modelled as an exponential CIR process.
problem Maximizing discounted spendings/dividend payments given an exponential CIR discounting factor.
method Analytical and numerical methods for deterministic and stochastic surplus processes.
result Explicit expressions for optimal strategies in deterministic cases, and constant-barrier strategies for small volatility in stochastic cases.
SA-BCP combines long-term and local evidence for efficient, adaptive online prediction.
problem Balancing fast adaptation and stable coverage in online prediction.
method State-Adaptive Bayesian Conformal Prediction (SA-BCP) using gated convex combination of temporal inertia and spatial evidence.
result SA-BCP achieves at-or-above-nominal coverage with substantially sharper intervals compared to discounted Bayesian CP.
New RL approach handles non-exponential discounting for sequential decisions.
problem Modeling human discounting in sequential decision-making tasks.
method Generalized model-based reinforcement learning with arbitrary discount functions, using Hamilton-Jacobi-Bellman equation and collocation method.
result Validated approach on simulated problems, showing applicability to human discounting.
The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …
We optimize discounts to maximize influence spread in social networks.
problem Maximizing influence spread in social networks with fractional discounts.
method Developed an efficient (1-1/e)-approximation algorithm for NP-hard problem.
result Achieved an approximation of 1-1/e for influence maximization.
Study optimal portfolio strategies with time-varying discount rates.
problem Optimizing portfolio decisions with a non-constant discount rate.
method Introduced subgame perfect strategies to handle time inconsistency, using fixed point iteration to find the utility-weighted discount rate.
result Subgame perfect strategies are equivalent to optimal strategies under certain utility function assumptions.
The study uses reproducing kernels to model bond discount curves.
problem Estimating bond discount curves under no-arbitrage conditions.
method Introduced reproducing kernels as a regression basis for estimating bond discount curves.
result Reproducing kernels provide a tractable solution for calibrating models to market data.
Policy gradient methods do not optimize the discounted objective, leading to suboptimal results.
problem Understanding the true optimization objective of policy gradient methods.
method Analyzing the update direction of policy gradient methods and proving it is not the gradient of any function.
result Policy gradient methods do not optimize the discounted objective, leading to suboptimal results.
We show that different rates should be used for borrowing and discount rates, and that the risk-free rate should be used for discounting when assessing and comparing the cost of energy accross diffferent producers and technologies, on the example of photovoltaics. Recent quantitative models using the same rate for borr…
New method uses entropy to improve policy gradient exploration.
problem Limited exploration in policy gradient methods.
method Entropy regularization with discounted future state distribution.
result Proves convergence to locally optimal policy.
We consider a discounted reward control problem in continuous time stochastic environment where the discount rate might be an unbounded function of the control process. We provide a set of general assumptions to ensure that there exists a smooth classical solution to the corresponding HJB equation. Moreover, some verif…
The paper proposes a method to discount backtest PnLs due to in-sample overfitting.
problem In-sample overfitting in backtest-based investment strategies.
method A simple framework to model and quantify in-sample PnL overfitting.
result Computes the appropriate discount factor for PnLs of in-sample investment strategies.
New findings reveal discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
problem Discount regularization leads to poor performance in unevenly sampled data.
method Equivalence theorem showing discount regularization as a strong prior, setting regularization parameters locally for individual state-action pairs.
result Discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
UCBVI-γ algorithm minimizes regret in discounted MDPs.
problem Minimizing regret in discounted MDPs.
method Optimism in the face of uncertainty principle and Bernstein-type bonus.
result UCBVI-γ achieves nearly minimax optimal regret.
New RL difficulty shown for discounted settings.
problem Difficulty in reinforcement learning with discounted rewards.
method Adapted Wang et al. (2020) construction to 2-state MDP.
result Learning impossible even with infinite data in discounted setting.
Lower discount factors act as a regularizer in RL, improving performance.
problem Improving RL performance with limited data.
method Explicitly equating reduced discount factors to regularization terms.
result Regularization effectiveness depends on data properties.
Optimality of threshold strategies proven for Lévy models with discounting.
problem Proving optimality of threshold strategies in Lévy models with discounting.
method Average problem approach to prove optimality of threshold strategies for Lévy models with continuous additive functional discounting.
result Simpler and neater proofs for qualitative properties of optimal thresholds in recursive optimal stopping problems.
Paper proposes a machine learning method to predict sale efficacy.
problem Determining the efficacy of online sales from discounts alone.
method Machine learning-based heuristic using Support Vector Machine.
result Predicts sale efficacy with 91.11% accuracy.
Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…
In this paper, we study the dividend strategies for a shareholder with non-constant discount rate in a diffusion risk model. We assume that the dividends can only be paid at a bounded rate and restrict ourselves to the Markov strategies. This is a time inconsistent control problem. The extended HJB equation is given an…
Dynamic promotion optimization for e-commerce platforms within financial constraints.
problem Balancing promotional costs with incremental revenue for sustainable growth.
method Knapsack Problem formulation for dynamic optimization, Retrospective Estimation, online-dynamic calibration.
result Significant increase in target outcome while staying within financial constraints.
Paper improves TD learning algorithm bounds with linear approx.
problem Sharp bounds for TD method performance in MDPs.
method Polyak-Ruppert averaging, universal step size, refined error bounds, stability of random matrices.
result Near-optimal variance and bias terms achieved.
A new method maps value estimates to logarithmic space to enable lower discount factors in reinforcement learning.
problem The poor performance of low discount factors in reinforcement learning.
method Introducing a logarithmic mapping to value estimates.
result The method enables lower discount factors, solving challenging reinforcement learning problems.