Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

21416282 · Jun 202019922001200920172026
48 results for Non-Markovian Rewards

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.

Enhances reward specification in RL with a novel language-based approach.

problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.

Efficient RL in PRMs with improved regret bound.

problem Reinforcement learning in probabilistic reward machines with non-Markovian rewards.
method Design of an algorithm with a new regret bound of O~(HOAT+H2O2A3/2+HT)\widetilde{O}(\sqrt{HOAT} + H^2O^2A^{3/2} + H\sqrt{T}).
result Improved regret bound over existing methods, matching lower bound up to a logarithmic factor.

This work analyzes QQ-learning with adaptive stepsizes for finite-time convergence.

problem Finite-time convergence analysis for average-reward QQ-learning with adaptive stepsizes.
method Adaptive stepsizes as local clocks, time-inhomogeneous Markovian reformulation, almost-sure time-varying bounds, conditioning arguments, and Markov chain concentration inequalities.
result Convergence rates of ildeO(1/k) ilde{\mathcal{O}}(1/k) for mean-square and pointwise mean-square convergence.

Paper introduces MVS to detect non-Markovian observations in reinforcement learning.

problem Real-world sensors violate Markov property, leading to suboptimal reinforcement learning performance.
method Uses prediction-based Markov Violation Score (MVS) combining random forest and ridge regression.
result MVS detects non-Markovian structure in observation trajectories, quantifying its impact.

Non-Markovian point process shows power-law scaling, similar to nonlinear Markovian process.

problem Understanding the scaling behavior of non-Markovian point processes.
method Analyzed a confined fractional Brownian motion-driven point process and compared it to a nonlinear Markovian process.
result A nonlinear Markovian process can reproduce the power-law scaling behavior of a non-Markovian point process.

In many settings (e.g., robotics) demonstrations provide a natural way to specify tasks; however, most methods for learning from demonstrations either do not provide guarantees that the artifacts learned for the tasks, such as rewards or policies, can be safely composed and/or do not explicitly capture history dependen…

2019-07-26abs ↗pdf ↗

Partially observable environments present an important open challenge in the domain of sequential control learning with delayed rewards. Despite numerous attempts during the two last decades, the majority of reinforcement learning algorithms and associated approximate models, applied to this context, still assume Marko…

2017-05-31abs ↗pdf ↗

Paper tackles robust offline RL for non-Markovian processes, improving efficiency and applicability.

problem Learning robust policies for non-Markovian decision processes with limited offline data.
method Proposes a novel algorithm with dataset distillation and LCB design for robust values, derived new dual forms, and introduces concentrability coefficients.
result Proves polynomial sample efficiency for finding ε-optimal robust policies.

Unified analytical tool for non-Markovian jump processes.

problem Analyzing history-dependent jump processes with non-Markovian behavior.
method Developed a standard form of master equations using Laplace-space embedding and asymptotic solution.
result Unified analytical toolset for general non-Markovian processes, leading to the GLE approximation.

Investigates optimal consumption and investment strategies in non-Markovian markets with unbounded parameters.

problem Optimal consumption and investment strategies in non-Markovian markets with unbounded parameters.
method Martingale optimal principle and quadratic BSDEs with exponential moment.
result Establishes optimal strategies for consumption and investment.

The paper develops a deep signature approach for option pricing under non-Markovian stochastic volatility models.

problem Pricing options under non-Markovian stochastic volatility models is challenging due to the dependence on historical paths.
method Reformulate the asset dynamics as a rough stochastic differential equation and represent rough paths via signatures. Apply standard analytical tools to solve the transformed equation.
result The deep signature approach provides a theoretically grounded and computationally efficient framework for option pricing.

Study optimal portfolios in a non-Markovian regime-switching model with random time horizon.

problem Optimal portfolio selection in a market with non-Markovian regime-switching and random time horizon.
method Formulated as a constrained stochastic linear-quadratic optimal control problem, derived closed-form expressions for optimal portfolios and efficient frontier.
result Closed-form expressions for optimal portfolios and efficient frontier derived under non-Markovian regime-switching and random time horizon.

Study develops numerical schemes for non-Markovian volatility models with memory.

problem Existence and uniqueness of strong solutions for non-Markovian SDEs.
method Functional quantization scheme based on Lamperti transformation.
result Theoretical foundation for numerical schemes applied to specific models.

Study analyzes non-Markovian effects in financial markets over multiple years.

problem Understanding non-Markovian dynamics and trader interactions in financial markets.
method Empirical analysis of self-response functions and trade sign correlators for different stocks over multiple years.
result Significant variations in traders' interactions over time, indicating changes in market mechanisms.

Paper tackles non-Markovian control problems with new learning methods.

problem Non-Markovian stochastic control problems with unknown parameters.
method Off-model training and importance sampling for deep neural network approximation.
result Quantitative error bounds for adaptive learning under model uncertainty.

This paper studies a class of non-Markovian singular stochastic control problems, for which we provide a novel probabilistic representation. The solution of such control problem is proved to identify with the solution of a ZZ-constrained BSDE, with dynamics associated to a non singular underlying forward process. Du…

2017-01-30abs ↗pdf ↗

Path signatures improve hedging of exotic derivatives in non-Markovian models.

problem Hedging exotic derivatives under non-Markovian stochastic volatility models.
method Investigates path signatures in deep and shallow learning contexts, comparing neural networks and regression approaches.
result Path signatures outperform LSTM in most cases and yield more accurate results in hedging.

The paper develops methods to price options under rough volatility models using BSPDEs.

problem Pricing options in models with non-Markovian dynamics.
method Backward stochastic partial differential equations (BSPDEs) and deep learning for numerical approximations.
result Existence and uniqueness of weak solutions for general nonlinear BSPDEs.

A new method predicts non-Markovian closure terms for complex systems.

problem Predicting the effect of unresolved variables on resolved dynamics in high-dimensional systems.
method Mamba-Assisted Closure (MAC) framework: sequence model trained to predict closure from resolved trajectory, coupled with reduced-order equations.
result Substantially outperforms existing methods in predictive accuracy and long-time stability.

HS-FNO models non-Markovian PDEs by learning history and future states.

problem Non-Markovian dynamics where future states depend on past history.
method History-Space Fourier Neural Operator (HS-FNO) for delay and memory-driven PDEs.
result HS-FNO achieves lowest aggregate errors across various PDE families.

New model controls memory in seq2seq tasks, revealing learning regimes.

problem Understanding memory in seq2seq tasks using neural networks.
method Introducing a stochastic switching-Ornstein-Uhlenbeck (SSOU) model to control memory and a measure of non-Markovianity.
result Two learning regimes emerge from the interplay of time scales in the SSOU process.

Representation learning on networks offers a powerful alternative to the oft painstaking process of manual feature engineering, and as a result, has enjoyed considerable success in recent years. However, all the existing representation learning methods are based on the first-order network (FON), that is, the network th…

2019-08-15abs ↗pdf ↗

Paper solves Merton's portfolio problem in a non-Markovian, non-semimartingale model.

problem Merton's portfolio optimization in a fake stationary Volterra-Heston model.
method Stochastic factor solution to a Riccati BSDE, combined with martingale optimality principle.
result Derives semi-closed form optimal strategies and value function.

Study on kinetic Langevin diffusions and their couplings, showing subtle TV bounds and new non-Markovian couplings.

problem Understanding and quantifying the TV distance between solutions of kinetic Langevin diffusions with different initial values.
method Established new non-Markovian couplings for kinetic Langevin diffusions, derived from optimal coalescence trajectories, and analyzed their TV bounds.
result No Markovian coupling can capture the asymptotic decay rate of the TV distance between solutions of kinetic Langevin diffusions with different initial values.

New approach tackles non-Markovian behavior in maternal health programs.

problem Improving adherence and engagement in maternal and child healthcare programs.
method Extending RMABs to non-Markovian settings, using time-series forecasting and TARI policy.
result Significant increase in engagement and content listened compared to existing methods.

FLDD improves discrete diffusion models by learning a non-Markovian noising process.

problem Efficiency and quality of discrete diffusion models in few-step generation.
method Introduces a learnable non-Markovian forward (noising) process to match the target distribution.
result FLDD produces higher quality samples in fewer steps compared to conventional discrete diffusion models.

Develops non-Markovian couplings for sub-Riemannian Brownian motions.

problem Constructing couplings for sub-Riemannian Brownian motions starting from points on the same vertical fiber.
method Uses global isometries to construct maximal couplings, satisfying a reflection principle.
result Estimates coupling time and applies to inequalities for the heat semigroup.

Study proves value of non-Markovian games with partial, asymmetric info.

problem Value of non-Markovian Dynkin games with partial and asymmetric information.
method Probabilistic and functional analytic approach based on Sion's min-max theorem.
result Existence of optimal strategies for both players in randomised stopping times.

Using a relationship between the moments of the probability distribution of times between the two consecutive trades (intertrade time distribution) and the moments of the distribution of a daily number of trades we show, that the underlying point process generating times of the trades is an essentially non-markovian lo…

2004-03-18abs ↗pdf ↗

Paper introduces IO-NPF for efficient Bayesian experimental design.

problem Efficient Bayesian experimental design in non-exchangeable settings.
method Inside-Out Nested Particle Filter (IO-NPF) for non-Markovian state-space models.
result IO-NPF achieves O(T2)\mathcal{O}(T^2) computational complexity, improving efficiency.

ARL bridges non-Markovian decision processes with reinforcement learning, improving foresight and stability.

problem Inaccurate foresight in non-Markovian environments due to state-based methods' limitations.
method Lifted state space into a signature-augmented manifold, using a self-consistent field approach to anticipate future path-law.
result ARL achieves deterministic evaluation of expected returns with reduced computational complexity and variance.

A new model predicts price concavity and reversion after metaorder execution.

problem Modeling market response to exogenous trades on limit order books.
method Developed a Non-Markovian Zero Intelligence model with a time-weighted mid-price return function.
result The model predicts concave price paths and price reversion after metaorder execution.

The paper values variable annuities using complex stochastic models and deep learning.

problem Valuation of variable annuities with early surrender options under non-Markovian models.
method Developed a deep signature Least Squares Monte Carlo approach to handle path-dependent continuation values.
result Fair fees increase with Hurst parameters of stock volatility and mortality force.