Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

131262393524 · Jun 202019922001200920172026
48 results for path dependent policies

Develops a new trading strategy for statistical arbitrage with path-dependent signals.

problem Optimal execution in statistical arbitrage strategies with dynamic predictive signals.
method Signature-based framework modeling alpha and trading speed as linear functionals of truncated signature of market path.
result Fitted policy achieves higher return on turnover compared to a z-score benchmark.

New algorithm achieves data-dependent regret bounds in MDPs with unknown transitions.

problem Achieving best-of-both-worlds guarantees with data-dependent regret bounds in MDPs with unknown transitions.
method Optimistic follow-the-regularized-leader algorithm with new optimistic Q-function estimators and transition bonus.
result First-order, second-order, and path-length bounds with polylog(T) regret in the stochastic regime.

New algorithms reduce regret in online MDPs by adapting to data and variance.

problem Adapting to both adversarial and stochastic environments in online MDPs.
method Develops algorithms based on global optimization and policy optimization, using optimistic follow-the-regularized-leader with log-barrier regularization.
result Achieves refined data-dependent and variance-dependent regret bounds.

We present the symmetric thermal optimal path (TOPS) method to determine the time-dependent lead-lag relationship between two stochastic time series. This novel version of the previously introduced TOP method alleviates some inconsistencies by imposing that the lead-lag relationship should be invariant with respect to …

2014-08-24abs ↗pdf ↗

ScoreMatchingRiesz improves debiased machine learning and policy effects estimation.

problem Improving debiased machine learning and policy effects estimation.
method Score matching and Riesz representer estimation.
result Estimates policy path for continuous treatments, improving interpretability.

We propose an algorithm for deterministic continuous Markov Decision Processes with sparse rewards that computes the optimal policy exactly with no dependency on the size of the state space. The algorithm has time complexity of O(R3×A2)O( |R|^3 \times |A|^2 ) and memory complexity of O(R×A)O( |R| \times |A| ), where R|R| is the…

2018-05-17abs ↗pdf ↗

Extend classical theory of affine processes to path-dependent setting

problem Path-dependent affine processes
method Introduce path-dependent coefficients and provide analytic formulas for their Fourier--Laplace transform
result Define path-dependent affine processes through their exponential-affine Fourier--Laplace transform and establish a characterization theorem

The paper develops methods to price and hedge options in path-dependent stock models.

problem Pricing and hedging options under complex stock models.
method Develops a path-dependent PDE for option pricing and differentiability of path-dependent SDE solutions.
result Provides formulas for option Greeks and differentiability of path-dependent SDE solutions.

We consider model-free reinforcement learning for infinite-horizon discounted Markov Decision Processes (MDPs) with a continuous state space and unknown transition kernel, when only a single sample path under an arbitrary policy of the system is available. We consider the Nearest Neighbor Q-Learning (NNQL) algorithm to…

2018-02-12abs ↗pdf ↗

This paper proposes a new approach to RL by focusing on the value-improvement path.

problem Value prediction problems in RL are sequence-dependent and require holistic approach.
method Characterize and approximate the value-improvement path holistically.
result A representation that spans the value-improvement path provides accurate value approximations for future policy improvements.

Study examines how economic policy uncertainty impacts stock markets.

problem Dynamic relationship between economic policy uncertainty and stock markets.
method Used symmetric thermal optimal path (TOPS) method.
result Different interaction patterns observed in emerging and developed markets.

DSPI connects natural policy gradient to policy iteration, proving global convergence.

problem Optimizing policies in reinforcement learning.
method DSPI framework, combining smoothed policy iteration and natural policy gradient.
result DSPI achieves geometric convergence and optimal complexity for policy optimization.

This paper improves tail dependence analysis by introducing a path-based approach.

problem The classical tail dependence coefficient fails to capture non-exchangeable features of tail dependence.
method The paper introduces a path-based maximal tail dependence approach to capture the most pronounced feature of dependence over all possible paths.
result The paper proves the existence and provides an explicit characterization of the path-based maximal TDC, improving analytical and computational tractability.

Develops neural network framework for risk-reward optimization problems.

problem Multi-period risk-reward optimization with constrained policies.
method Neural network framework with two coupled feedforward networks, parametrizing two-step policies.
result Empirical optimum converges to true optimal value as network capacity and training size increase.

New algorithm reduces regret in stochastic shortest path problems.

problem Planning and control in environments with unknown dynamics and variable episode lengths.
method Developed an algorithm with a new regret bound of O(BSAK)O(B_\star |S| \sqrt{|A| K}).
result Guaranteed a significant reduction in regret compared to previous methods.

Develops a numerical scheme for solving path-dependent FBSDEs and PDEs.

problem Solving path-dependent FBSDEs and PDEs numerically.
method Picard iteration method for FBSDEs, concentration inequality for estimator, supervised learning with neural networks for PDEs.
result Proves convergence and rate of convergence for the Picard iteration method.

Dupire's functional Itô calculus provides an alternative approach to the classical Malliavin calculus for the computation of sensitivities, also called Greeks, of path-dependent derivatives prices. In this paper, we introduce a measure of path-dependence of functionals within the functional Itô calculus framework. Name…

2013-11-15abs ↗pdf ↗

The study examines insurance demand under rough volatility and path-dependent shocks.

problem Optimal insurance and investment strategies under rough volatility and path-dependent shocks.
method Rough volatility model and Hawkes process with power kernel, Functional Ito formula extension.
result Individuals demand more catastrophe insurance when path-dependent effects are considered.

Deep signature algorithm for pricing path-dependent options.

problem Pricing path-dependent options with complex payoff functions.
method Extended backward scheme for state-dependent FBSDEs with reflections, incorporating signature layer for path-dependent FBSDEs.
result Convergence analysis of the algorithm with explicit dependence on truncation order and neural network approximation errors.

Study shows sample complexity for learning optimal policies in SSP with generative model.

problem Learning optimal policies in Stochastic Shortest Path problems.
method Derive and prove lower and upper bounds on sample complexity.
result Lower bound of Ω(SAB3/(cminε2))Ω(SAB_{\star}^3/(c_{\min}ε^2)) samples for general case, and up to logarithmic factors for bounded hitting time condition.

Study shows hard sample complexity for learning optimal policies in stochastic shortest path problems.

problem Learning optimal policies in stochastic shortest path problems.
method Analyzes sample complexity with and without generative models, derives lower and upper bounds.
result Proves sample complexity bounds and impossibility of horizon-free regret in SSPs.

Recent years have witnessed a tremendous improvement of deep reinforcement learning. However, a challenging problem is that an agent may suffer from inefficient exploration, particularly for on-policy methods. Previous exploration methods either rely on complex structure to estimate the novelty of states, or incur sens…

2019-11-11abs ↗pdf ↗

Path-dependent PDEs model VIX and Realised Variance options.

problem Modeling volatility derivatives with path-dependence.
method Continuous stochastic volatility model with Gaussian Volterra process, proving well-posedness of PDEs.
result Formulae for greeks and implied volatility provided, finite-dimensional pricing PDEs obtained in Markovian models.

ARL bridges non-Markovian decision processes with reinforcement learning, improving foresight and stability.

problem Inaccurate foresight in non-Markovian environments due to state-based methods' limitations.
method Lifted state space into a signature-augmented manifold, using a self-consistent field approach to anticipate future path-law.
result ARL achieves deterministic evaluation of expected returns with reduced computational complexity and variance.

Paper proposes method for generating paths of stochastic volatility CGMY process for option pricing.

problem Generating accurate sample paths for stochastic volatility models for option pricing.
method Monte-Carlo method for European and American options, least square regression for calibration.
result Calibrated model parameters to S\&P 100 index options market using path-dependent options.

Deep neural networks solve stochastic control problems with delay.

problem Challenges in stochastic control problems with delay due to path-dependence and high dimensions.
method Employing recurrent neural networks (RNNs) to parameterize policies and optimize objectives.
result RNNs, especially LSTMs, efficiently capture path-dependence and outperform feedforward networks in training and performance.

The paper provides an efficient method to price path-dependent derivatives using multiscale stochastic volatility models.

problem Pricing path-dependent derivatives under multiscale stochastic volatility models.
method Derives a Malliavin representation for the first-order approximation of the price of path-dependent derivatives.
result An efficient Monte Carlo approximation for pricing path-dependent derivatives is derived.

Extends Itô's formula for path-dependent functions in finance.

problem Modeling and hedging of path-dependent financial options.
method Functional extension of Itô's formula for C^{0,1}-functions of continuous weak Dirichlet processes.
result Validates the hedging or superhedging problems for path-dependent options.

Path integral method calculates PDBS option prices with time-dependent parameters.

problem Pricing proportional double-barrier step options with time-dependent interest rates and volatilities.
method Path integral method applied to a quantum mechanical analogy of barrier options.
result Derivation of pricing kernel for PDBS options with time-dependent parameters.

The paper develops a method to learn cost-optimal sequential testing policies from retrospective data.

problem Learning cost-optimal sequential decision policies from retrospective data with missing test results.
method Doubly robust Q-learning framework with path-specific inverse probability weights.
result The method reduces testing cost without compromising predictive accuracy.

PDGM uses neural nets to solve complex financial equations.

problem Solving path-dependent partial differential equations (PPDEs)
method Generalized Deep Galerkin Method (PDGM) combining feed-forward and LSTM architectures
result PDGM successfully models solutions to various PPDEs, including financial derivatives.

New control theory for self-path-dependent problems solves unique constraints.

problem Optimal control with self-path-dependent constraints in stochastic systems.
method Introduces new HJB equations for variational inequalities with historical maximum controls.
result Value functions are viscosity solutions to HJB equations under Lipschitz conditions.