Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

7.8%15.6%23.4%31.2% · Jun 202019922001200920182026
48 results for Value-function Optimization

Geometric approach improves reinforcement learning representation.

problem Improving reinforcement learning representation learning.
method Formal evidence through geometric properties of value functions.
result Optimizing value functions reduces to predicting adversarial value functions (AVFs).

POLO framework enables efficient learning and exploration in model-based control.

problem Efficient learning and exploration in model-based control settings.
method Combines local model-based control, global value function learning, and exploration.
result POLO framework accelerates value function learning and enables better policies.

This paper interpolates reward functions to predict optimal value functions in MORL.

problem Finding optimal value functions in MORL requires recomputing for each set of weights.
method Interpolating reward function weights to smooth value function transformations.
result Smooth interpolation of optimal value functions over reward function weights.

A new method separates long-term value functions into components for better reinforcement learning.

problem Learning long-term goals in reinforcement learning settings with temporal discounting.
method TD(ΔΔ) learning, which breaks down value functions into components based on discount factors.
result TD(ΔΔ) learning improves scalability and performance over standard TD learning in certain settings.

Paper introduces a new value function for state transitions and optimal policy learning.

problem Learning optimal policies from state transitions and actions.
method Develops a forward dynamics model to maximize a novel value function Q(s,s)Q(s, s').
result Demonstrates benefits in value function transfer, redundant action spaces, and off-policy learning.

Develops a new method for statistical optimal allocation problems.

problem Statistical optimal allocation problems with constraints.
method Functional differentiability approach and Hadamard differentiability of value functions.
result Validates margin assumption for fast convergence rate of plug-in methods.

This paper considers a utility maximization and optimal asset allocation problem in the presence of a stochastic endowment that cannot be fully hedged through trading in the financial market. After studying continuity properties of the value function for general utility functions, we rely on the dynamic programming app…

2014-06-24abs ↗pdf ↗

PPG separates policy and value function training phases for better reinforcement learning efficiency.

problem Challenges in traditional reinforcement learning methods for policy and value function optimization.
method Integrates Phasic Policy Gradient framework that splits policy and value function training into distinct phases.
result Significantly improves sample efficiency on Procgen Benchmark compared to PPO.

Optimizes dividend payout in insurance wealth process with stochastic interest rate.

problem Maximizing expected discounted dividends up to ruin in insurance wealth process.
method Modelled compound Poisson process with stochastic interest rate, solved using HJB equation.
result Explicit expression for value function and optimal strategy in geometric Brownian motion case.

Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true value function VV. Two novel algorithms are proposed to approximate the true value…

2017-04-17abs ↗pdf ↗

The paper examines smoothness of value function in consumption-investment models with borrowing constraints.

problem Investor's optimal consumption and investment under consumption-wealth utility and borrowing constraint.
method Second-order smoothness of value function, optimal consumption-investment policy in feedback form, smooth fit condition.
result The value function is second-order smooth and the constraint is binding under certain conditions.

TensorPlan algorithm finds δ-optimal policies with poly(H,d)(H,d) queries under linearly realizable state-value function.

problem Efficient planning in MDPs with linearly realizable state-value function.
method TensorPlan algorithm using poly((dH/δ)A)((dH/δ)^A) simulator queries.
result First algorithm with polynomial query complexity using only linear-realizability of a single competing value function.

Addressing RL's agent-environment boundary issues, a novel analysis ensures optimal value functions are invariant.

problem Fundamental RL concepts like value functions are not uniquely defined due to the agent-environment boundary.
method A boundary-invariant analysis of Fitted Q-Iteration, ensuring optimality guarantees are independent of the boundary choice.
result Theoretical analyses of RL algorithms, including Fitted Q-Iteration, are made invariant to the boundary choice.

New concept of Blackwell regret for reinforcement learning with sparse rewards.

problem Sparse rewards in long horizon MDPs.
method Formalization of myopic discount factors, value functions, and policies in terms of Blackwell optimality; introduction of Blackwell regret.
result Selecting a discount factor for zero Blackwell regret becomes arbitrarily hard in long horizon MDPs.

The Lax-Hopf formula simplifies the value function of an intertemporal optimization (infinite dimensional) problem associated with a convex transaction-cost function which depends only on the transactions (velocities) of a commodity evolution: it states that the value function is equal to the marginal fonction of a fin…

2014-01-08abs ↗pdf ↗

Optimizes dividend and reinsurance strategies for correlated insurance lines.

problem Stochastic control of optimal reinsurance and dividend policies for multiple insurance lines.
method Maximizes cumulative discounted dividends using a Hamilton-Jacobi-Bellman equation and finite difference method.
result Provides optimal strategies for transferring risk among reinsurers.

Develops robust MDPs for unknown disturbances with performance guarantees.

problem Unknown disturbance distribution in MDPs.
method Empirical distribution, sublevel set of distance function, weak convergence, concentration inequality.
result Robust optimal value function converges to true optimal value function with increasing sample sizes.

We study an optimal execution problem in a continuous-time market model that considers market impact. We formulate the problem as a stochastic control problem and investigate properties of the corresponding value function. We find that right-continuity at the time origin is associated with the strength of market impact…

2009-07-20abs ↗pdf ↗

Paper approximates free boundary for optimal investment stopping problems.

problem Optimal investment stopping problems with utility maximization.
method Dual control method to derive asymptotic properties and construct a global closed-form approximation.
result Global closed-form approximation of dual free boundary reduces computational cost.

Efficiently plans large MDPs with weak function approximations.

problem Planning in large MDPs with limited function approximation capabilities.
method Uses linear value function approximation with weak requirements and a generative oracle.
result Produces almost-optimal actions for any state with polynomial computation time.

We consider the classical optimal dividends problem under the Cramér-Lundberg model with exponential claim sizes subject to a constraint on the time of ruin. We introduce the dual problem and show that the complementary slackness conditions are satisfied, thus there is no duality gap. Therefore the optimal value functi…

2014-10-14abs ↗pdf ↗

We characterize value functions in partially observable MDPs as semi-algebraic sets.

problem Understanding feasible value functions in partially observable Markov decision processes.
method Characterization of feasible value functions as semi-algebraic sets defined by polynomial inequalities.
result The feasible set of value functions in POMDPs is a semi-algebraic set, not a polytope as in MDPs.

Paper tackles robust control of SDEs with ambiguity, proving value function existence and applying to investment problems.

problem Robust control of SDEs with ambiguity parameters and non-Lipschitz coefficients.
method Existence and uniqueness of value function established through BSDEs with non-linear growth conditions.
result Existence and uniqueness of value function in proper space, verified through BSDEs.

Estimates personalized treatment response curves using covariates.

problem Flexible estimation of personalized treatment response curves.
method Sieve based nonparametric estimator of smoothed regimen-response curve function.
result Asymptotic linearity and undersmoothing criteria for efficient estimation.

The value function of an optimal stopping problem for jump diffusions is known to be a generalized solution of a variational inequality. Assuming that the diffusion component of the process is nondegenerate and a mild assumption on the singularity of the Lévy measure, this paper shows that the value function of this op…

2009-02-15abs ↗pdf ↗

We study the optimal liquidation problem in a market model where the bid price follows a geometric pure jump process whose local characteristics are driven by an unobservable finite-state Markov chain and by the liquidation rate. This model is consistent with stylized facts of high frequency data such as the discrete n…

2016-06-16abs ↗pdf ↗

Study optimal liquidation strategies with infinite horizon and regime switching.

problem Optimal liquidation with semimartingale strategies in a stochastic environment.
method Characterization of value function and optimal strategy via BSDEs with infinite horizon.
result Existence and uniqueness of optimal control problem solutions.

We consider the problem of portfolio optimization in a simple incomplete market and under a general utility function. By working with the associated Hamilton-Jacobi-Bellman partial differential equation (HJB PDE), we obtain a closed-form formula for a trading strategy which approximates the optimal trading strategy whe…

2016-11-28abs ↗pdf ↗

Optimizing dividend payments for an insurance company with bounded rates.

problem Maximizing expected exponential utility of discounted dividends under bounded dividend rates.
method Suboptimal strategies are evaluated using a new method to estimate the distance to the value function.
result The optimal strategy is of barrier type with a non-linear barrier.

We develop a polynomial method to optimize trading in markets with transaction costs.

problem Optimizing trading strategies in markets with proportional transaction costs.
method Polynomial approximation of the residual value function to determine optimal trading strategies.
result Identify the trade-off between trading frequency and trade sizes for satisfactory agreement with theoretically optimal strategies.