Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

70140209279 · Jun 202019922001200920172026
48 results for discounted objective

Policy gradient methods do not optimize the discounted objective, leading to suboptimal results.

problem Understanding the true optimization objective of policy gradient methods.
method Analyzing the update direction of policy gradient methods and proving it is not the gradient of any function.
result Policy gradient methods do not optimize the discounted objective, leading to suboptimal results.

New Q-learning algorithm reduces sample complexity for large discount factors.

problem Large discount factors make Q-learning algorithms inefficient.
method Introduces a new Q-learning algorithm with uniformly bounded sample complexity.
result The new algorithm achieves asymptotic covariance that is a quadratic in 1/(1ργ)1/(1- ρ^* γ).

The paper explores a Multi-Objective RL approach for trading that generalizes reward functions.

problem Improving performance in single-asset trading through adaptive reward functions.
method Developed a Multi-Objective Deep Reinforcement Learning algorithm to generalize reward functions and discount factors.
result The Multi-Objective algorithm demonstrates increased predictive stability and better performance in sparse reward scenarios.

The objective of the present paper is to analyse various features of the Smith-Wilson method used for discounting under the EU regulation Solvency II, with special attention to hedging. In particular, we show that all key rate duration hedges of liabilities beyond the Last Liquid Point will be peculiar. Moreover, we sh…

2016-02-05abs ↗pdf ↗

This paper presents an algorithm for pricing perpetual American put options with asset-dependent discounting.

problem Pricing perpetual American put options with asset-dependent discounting.
method The approach involves a value function described by a stochastic process with negative exponential jumps and a discount function that depends on the asset price.
result Under certain conditions, the value function can be convex and represented in a closed form.

The standard reinforcement learning (RL) formulation considers the expectation of the (discounted) cumulative reward. This is limiting in applications where we are concerned with not only the expected performance, but also the distribution of the performance. In this paper, we introduce micro-objective reinforcement le…

2019-05-24abs ↗pdf ↗

Paper develops a discounted algorithm for online convex optimization that adapts to unknown discount factors.

problem Developing an algorithm that can adapt to an unknown discount factor in online convex optimization.
method Smoothed Online Gradient Descent (SOGD) with Discounted-Normal-Predictor (DNP).
result Achieves a uniform O(logT/1λ)O(\sqrt{\log T/1-λ}) discounted regret across a continuous interval of discount factors.

Study optimal stopping times under regime-switching models with constraints.

problem Optimal stopping times for discounted payoffs on a regime-switching geometric Brownian motion.
method Solve variational inequality to find value functions and optimal thresholds.
result Existence and expressions of optimal stopping times under specific conditions.

Paper introduces non-linear discounting models for default compensation and climate valuation.

problem Valuation of non-replicable value and damage under default risk.
method Develops two models: one for risk-neutralising discounting and another for survival probability dependent discounting.
result Non-decaying discount factors (negative discount rates) are possible under certain scenarios.

New objective reduces bias and variance in reinforcement learning derivatives.

problem Estimating derivatives in reinforcement learning with unknown dynamics.
method Derives an objective function compatible with any advantage estimators, allowing trade-off between bias and variance.
result Demonstrates effectiveness in both theoretical and practical settings.

The Epicurean Philosophy is commonly thought as simplistic and hedonistic. Here I discuss how this is a misconception and explore its link to Reinforcement Learning. Based on the letters of Epicurus, I construct an objective function for hedonism which turns out to be equivalent of the Reinforcement Learning objective …

2017-10-12abs ↗pdf ↗

Study analyzes how discounts affect train ticket purchases and rescheduling in Switzerland.

problem Understanding how discounts influence train ticket buying and rescheduling behavior.
method Machine learning techniques, including causal machine learning, to analyze survey data.
result Increasing a discount rate by 1% increases the rescheduled trip share by 0.16% among always buyers.

Develops a machine-learning framework for optimal share repurchase hedging.

problem Challenges in hedging share repurchase programs due to market regulations and trading activity.
method Machine-learning framework that optimizes execution and hedging of share repurchase programs.
result Substantial performance improvements and an optimized hedging approach.

Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process. The discount factor values future rewards by an exponential scheme that leads to theoretical convergence guarantees of the Bellman equation. However, evidence from psychology, economics and neuroscience suggests that…

2019-02-19abs ↗pdf ↗

There is an observed basis between repo discounting, implied from market repo rates, and bond discounting, stripped from the market prices of the underlying bonds. Here, this basis is explained as a convexity effect arising from the decorrelation between the discount rates for derivatives and bonds. Using a Hull-White …

2019-05-08abs ↗pdf ↗

This paper shows how forward rate interpolations are equivalent to discount factor interpolations in yield curve construction.

problem The challenge of choosing between different interpolation methods for yield curve construction.
method Demonstrates the equivalence between forward rate interpolations and discount factor interpolations.
result Some popular interpolation methods on forward rates are equivalent to classical interpolation methods on discount factors.

New RL approach handles non-exponential discounting for sequential decisions.

problem Modeling human discounting in sequential decision-making tasks.
method Generalized model-based reinforcement learning with arbitrary discount functions, using Hamilton-Jacobi-Bellman equation and collocation method.
result Validated approach on simulated problems, showing applicability to human discounting.

Policy gradient methods with aggregated states can achieve better performance than approximate policy iteration.

problem Approximation errors in policy and value function approximations.
method State-aggregated representations and policy gradient methods.
result Policy gradient methods can achieve a per-period regret bounded by ε, while approximate policy iteration and value iteration have a higher regret.

The paper introduces a new intrinsic reward method for exploration in reinforcement learning.

problem Improving exploration in reinforcement learning agents.
method Intrinsic rewards proportional to the entropy of future state-action features.
result The new objective leads to improved visitation of features within individual trajectories.

The valuation process that economic agents undergo for investments with uncertain payoff typically depends on their statistical views on possible future outcomes, their attitudes toward risk, and, of course, the payoff structure itself. Yields vary across different investment opportunities and their interrelations are …

2010-01-08abs ↗pdf ↗

We optimize discounts to maximize influence spread in social networks.

problem Maximizing influence spread in social networks with fractional discounts.
method Developed an efficient (1-1/e)-approximation algorithm for NP-hard problem.
result Achieved an approximation of 1-1/e for influence maximization.

Study optimal portfolio strategies with time-varying discount rates.

problem Optimizing portfolio decisions with a non-constant discount rate.
method Introduced subgame perfect strategies to handle time inconsistency, using fixed point iteration to find the utility-weighted discount rate.
result Subgame perfect strategies are equivalent to optimal strategies under certain utility function assumptions.

Paper analyzes robust strategies in a pension plan game with ambiguous financial markets.

problem Analyzing robust strategies in a defined benefit pension plan game with ambiguous financial markets.
method Formulated and solved two robust non-zero-sum games using stochastic dynamic programming.
result Explicit forms and optimality of the solutions are shown for the firm and union.

New findings reveal discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.

problem Discount regularization leads to poor performance in unevenly sampled data.
method Equivalence theorem showing discount regularization as a strong prior, setting regularization parameters locally for individual state-action pairs.
result Discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.

Asset prices contain information about the probability distribution of future states and the stochastic discounting of those states as used by investors. To better understand the challenge in distinguishing investors' beliefs from risk-adjusted discounting, we use Perron-Frobenius Theory to isolate a positive martingal…

2014-11-28abs ↗pdf ↗

In this paper, we study the dividend strategies for a shareholder with non-constant discount rate in a diffusion risk model. We assume that the dividends can only be paid at a bounded rate and restrict ourselves to the Markov strategies. This is a time inconsistent control problem. The extended HJB equation is given an…

2013-04-30abs ↗pdf ↗

A new method maps value estimates to logarithmic space to enable lower discount factors in reinforcement learning.

problem The poor performance of low discount factors in reinforcement learning.
method Introducing a logarithmic mapping to value estimates.
result The method enables lower discount factors, solving challenging reinforcement learning problems.

We examine two different techniques for parameter averaging in GAN training. Moving Average (MA) computes the time-average of parameters, whereas Exponential Moving Average (EMA) computes an exponentially discounted sum. Whilst MA is known to lead to convergence in bilinear settings, we provide the -- to our knowledge …

2018-06-12abs ↗pdf ↗