Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

165330495660 · Jun 202019922001200920182026
48 results for Semi Markov Decision Processes

We characterize value functions in partially observable MDPs as semi-algebraic sets.

problem Understanding feasible value functions in partially observable Markov decision processes.
method Characterization of feasible value functions as semi-algebraic sets defined by polynomial inequalities.
result The feasible set of value functions in POMDPs is a semi-algebraic set, not a polytope as in MDPs.

Develops new reinforcement learning methods for complex constrained decision-making problems.

problem Complex constrained decision-making problems with a continuum of constraints.
method Proposes semi-infinitely constrained Markov decision processes (SICMDPs) and two reinforcement learning algorithms: SI-CRL and SI-CPO.
result Demonstrates the effectiveness of SI-CRL and SI-CPO in solving complex sequential decision-making tasks.

The paper tackles robust policy learning in MDPs using statistical methods.

problem Offline data-driven sequential decision making in MDPs.
method Evaluates policies using average rewards centered at policy-induced stationary distributions. Developed a statistically efficient method for estimating robust optimal policies.
result Established a rate-optimal regret bound up to a logarithmic factor.

Automatically discovers SMDP models in DQN representations for better reinforcement learning visualization.

problem Lack of tools to analyze and visualize the temporal abstractions learned by DRL agents.
method Develops a novel method to automatically discover an internal SMDP model in DQN representations and visualizes it using a directed graph above a t-SNE map.
result Shows evidence of hierarchical state aggregation learned by DQNs.

In this paper we propose a semi-Markov modulated model of interest rates. We assume that the switching process is a semi-Markov process with finite state space E and the modulated process is a diffusive process. We derive recursive equations for the higher order moments of the discount factor and we describe a Monte Ca…

2012-10-11abs ↗pdf ↗

Paper proposes a DRL-based controller for networked AP systems that reduces communication frequency.

problem Reduce communication frequency in networked AP systems while maintaining control performance.
method Develops a DRL-based controller that avoids explicit update timing learning, using a semi-Markov decision process (SMDP).
result Improves communication efficiency without sacrificing control performance.

New approach uses deep reinforcement learning for vehicle dispatching, reducing waiting times.

problem Dynamic vehicle dispatching problem in various contexts.
method Event-based semi-Markov decision process with deep q-learning.
result Deep reinforcement learning policies outperform heuristic methods in New York City data.

Paper tests Markov assumption in sequential decision making.

problem Testing the Markov assumption in sequential decision making.
method Forward-Backward Learning procedure to test MA without assuming parametric forms.
result The proposed test plays a crucial role in identifying optimal policies in complex decision processes.

New algorithm solves uncertain Markov decision processes using Wasserstein uncertainty.

problem Solving Markov decision processes with uncertain transition probabilities.
method Distributionally robust QQ-learning algorithm for Wasserstein uncertainty.
result Convergence of the algorithm proved and demonstrated with real data.

A new method for semi-supervised classification using graph walks and reinforcement learning.

problem Efficiently classifying nodes in attributed networks with limited labeled data.
method Proposes a reinforcement learning approach to find optimal paths in the graph for classification.
result The method outperforms existing approaches on multiple datasets.

The study extends asset pricing models to include time-dependent volatility and age-dependent regime switching.

problem Asset pricing in a market with time-varying interest rates and volatilities.
method Extension of Markov-modulated models to semi-Markov processes with age-dependent and time-dependent volatility.
result Option pricing in the extended model is equivalent to solving an integral equation.

Study long-term behavior of semi-Markov modulated processes using integral functions.

problem Analyzing long-term behavior of semi-Markov modulated processes involving integral functions.
method Using ergodic semi-Markovian environment and affine stochastic recurrence equation.
result Mixture type laws emerge in long-term limit for processes.

A method for low-dimensional MDP representation using deep neural networks.

problem Constructing a low-dimensional Markov decision process representation.
method Use a deep neural network to define a class of potential process representations and estimate the process of lowest dimension.
result A decision strategy that maximizes mean utility for the low-dimensional representation also maximizes mean utility for the original process.

We study the high frequency price dynamics of traded stocks by a model of returns using a semi-Markov approach. More precisely we assume that the intraday return are described by a discrete time homogeneous semi-Markov process and the overnight returns are modeled by a Markov chain. Based on this assumptions we derived…

2011-03-31abs ↗pdf ↗

Paper presents an algorithm for optimal regret in communicating Markov decision processes.

problem Achieving optimal regret in Markov decision processes with a communicating assumption.
method The algorithm explicitly tracks the constant K(M) to learn optimally, balancing exploration, co-exploration, and exploitation.
result The algorithm achieves asymptotically optimal regret K(M)log(T)+o(log(T))K(M) \log(T) + \mathrm{o}(\log(T)) for communicating Markov decision processes.

Study optimality in safety-constrained Markov decision processes using asynchronous value iteration and modified Q-learning.

problem Optimality in safety-constrained Markov decision processes with multichain structure.
method Formulated as a zero-sum game, constructed asynchronous value iteration scheme and modified Q-learning algorithm.
result Resolved Bellman's principle of optimality for multichain Markov decision processes and provided learning algorithms.

New algorithms learn in complex decision-making problems with smooth transitions.

problem Learning in complex decision-making problems with smooth transitions.
method UCB and PSRL philosophies applied to episodic Markov decision processes with kernel approximation.
result Low regret learning achieved in continuous state and action spaces.

We consider the problem of constructing an appropriate multivariate model for the study of the counterparty credit risk in credit rating migration problem. For this financial problem different multivariate Markov chain models were proposed. However the markovian assumption may be inappropriate for the study of the dyna…

2011-12-01abs ↗pdf ↗

Overview of risk-sensitive Markov decision processes with Optimized Certainty Equivalent.

problem Optimizing decision-making under risk in Markov processes.
method Analyzes risk-sensitive criteria using Optimized Certainty Equivalent, including entropic risk and Conditional Value-at-Risk.
result Conditions for the existence of optimal policies and solution procedures are provided.

There is much interest in the Hierarchical Dirichlet Process Hidden Markov Model (HDP-HMM) as a natural Bayesian nonparametric extension of the ubiquitous Hidden Markov Model for learning from sequential and time-series data. However, in many settings the HDP-HMM's strict Markovian constraints are undesirable, particul…

2012-03-07abs ↗pdf ↗

New self-exciting random evolutions (SEREs) for modeling traffic and transport processes.

problem Modeling self-exciting and clustering effects in traffic and transport processes.
method Introducing a new process based on a superposition of a Markov chain and a Hawkes process, and constructing self-exciting random evolutions (SEREs).
result Developed new models and limit theorems for SEREs, including averaging and diffusion approximation.

Model stock price dynamics using semi-Markov processes.

problem Model stock price dynamics through a semi-Markov process.
method Use semi-Markov process with Poisson random measure, establish existence and uniqueness of solution, derive HJB equation.
result Obtain expressions for optimal controls and value function using HJB equation.

New method reduces sample complexity for robust reinforcement learning.

problem Finite sample analysis in robust reinforcement learning.
method Stochastic approximation framework with controlled bias, using MLMC techniques and geometric truncation.
result Order-optimal sample complexity of ildeO(ε2) ilde{\mathcal{O}}(ε^{-2}) for robust policy evaluation.

We optimize saddle-point problems for large-scale Markov decision processes.

problem Optimizing policies in large-scale Markov decision processes.
method Characterized conditions for convergence and designed an optimization algorithm.
result Our algorithm converges faster and is state-space independent.

Generalized model for firm valuation considering semi-Markovian dividend growth.

problem Valuation of firms based on semi-Markovian dividend growth rates.
method Discrete time semi-Markov chain model with measurable space, new equations for price-dividend ratios, approximation methods.
result Established sufficient conditions for finiteness of fundamental prices and risks, new equations for first and second order price-dividend ratios.

Algorithm finds safe zones in policy Markov Decision Processes to limit trajectory escape.

problem Finding safe zones in policy Markov Decision Processes to limit trajectory escape.
method Bi-criteria approximation learning algorithm with polynomial sample complexity.
result Achieves almost 2 approximation for both escape probability and safe zone size.

In this paper we propose a bivariate generalization of a weighted indexed semi-Markov chains to study the high frequency price dynamics of traded stocks. We assume that financial returns are described by a weighted indexed semi-Markov chain model. We show, through Monte Carlo simulations, that the model is able to repr…

2013-05-02abs ↗pdf ↗

Reinforcement learning struggles with corrupted reward signals.

problem Agents observe rewards that are not accurate due to sensory errors or software bugs.
method Formalized as Corrupt Reward MDP, investigated two approaches: richer data and randomisation.
result Traditional RL methods fail in CRMDPs, necessitating new approaches.

Risk measures applied to dynamic Markov processes with varying risk aversion.

problem Investigating dynamic risk measures in Markov decision processes with varying risk aversion.
method Distributional viewpoint on law-invariant convex risk measures, applied to Markov decision processes with latent costs and random actions.
result Existence of optimal policies in finite and infinite time horizons under mild assumptions.

Bayesian method infers local rules for collective animal movement.

problem Learn local rules governing long-term group behaviors.
method Bayesian Inverse Reinforcement Learning with Linearly-Solvable Markov Decision Process.
result Recover true costs and find value of collective movement.

The paper tackles batch policy learning in Markov Decision Processes, focusing on average reward maximization.

problem Maximizing long-term average reward in Markov Decision Processes with batch learning.
method Doubly robust estimator for average reward, optimization algorithm for optimal policy, finite-sample regret guarantee.
result The proposed method achieves semiparametric efficiency and provides a finite-sample regret guarantee.

Paper proposes a new method to model event sequences in information systems.

problem Analyzing event logs to understand system procedures and predict changes.
method Combines hidden semi-Markov model and classification trees learning.
result The proposed approach can identify frequent sequence patterns relevant to observable events.

Paper proposes an HMM-based Q-learning for POMDPs.

problem Q-learning struggles with POMDPs due to incomplete state observation.
method Formulates POMDP estimation as HMM estimation, proposing a recursive algorithm to concurrently estimate POMDP parameters and Q function.
result Algorithm converges to optimal Q function and POMDP parameters.

We study the high frequency price dynamics of traded stocks by a model of returns using a semi-Markov approach. More precisely we assume that the intraday returns are described by a discrete time homogeneous semi-Markov which depends also on a memory index. The index is introduced to take into account periods of high a…

2011-09-20abs ↗pdf ↗