Improved analysis of UCBVI algorithm with better empirical performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study an exploration method for model-free RL that generalizes the counter-based exploration bonus methods and takes into account long term exploratory value of actions rather than a single step look-ahead. We propose a model-free RL method that modifies Delayed Q-learning and utilizes the long-term exploration bonu…
Study minimax optimal RL in factored MDPs with bonus exploration.
The paper deals with bonus-malus systems with different claim types and varying deductibles. The premium relativities are softened for the policyholders who are in the malus zone and these policyholders are subject to per claim deductibles depending on their levels in the bonus-malus scale and the types of the reported…
New algorithm reduces reinforcement learning complexity, approaching contextual bandits.
The paper calculates bonus values in complex insurance schemes.
EQO uses a simple bonus term for efficient exploration in tabular RL.
Develops a Bonus-Malus model for cyber risk insurance to incentivize cybersecurity.
We discuss the pricing methodology for Bonus Certificates and Barrier Reverse-Convertible Structured Products. Pricing for a European barrier condition is straightforward for products of both types and depends on an efficient interpolation of observed market option pricing. Pricing products We discuss the pricing metho…
We introduce an exploration bonus for deep reinforcement learning methods that is easy to implement and adds minimal overhead to the computation performed. The bonus is the error of a neural network predicting features of the observations given by a fixed randomly initialized neural network. We also introduce a method …
Rewards are sparse in the real world and most of today's reinforcement learning algorithms struggle with such sparsity. One solution to this problem is to allow the agent to create rewards for itself - thus making rewards dense and more suitable for learning. In particular, inspired by curious behaviour in animals, obs…
The study analyzes how bonus-malus systems and delayed claims settlement affect insurance companies' financial stability.
We obtain Harnack estimates for a class of curvature flows in Riemannian manifolds of constant non-negative sectional curvature as well as in the Lorentzian Minkowski and de Sitter spaces. Furthermore, we prove a Harnack estimate with a bonus term for mean curvature flow in locally symmetric Riemannian Einstein manifol…
We introduce and analyse two algorithms for exploration-exploitation in discrete and continuous Markov Decision Processes (MDPs) based on exploration bonuses. SCAL is a variant of SCAL (Fruit et al., 2018) that performs efficient exploration-exploitation in any unknown weakly-communicating MDP for which an upper bo…
The paper introduces Bellman-consistent pessimism to improve offline reinforcement learning without overly pessimistic bias.
Paper analyzes strategic underreporting in competitive insurance markets.
New UCB algorithm for learning PSRs with tractable computation and accuracy.
The study finds solar terms significantly impact China's stock market returns and volatility.
In this work, we consider the popular tree-based search strategy within the framework of reinforcement learning, the Monte Carlo Tree Search (MCTS), in the context of infinite-horizon discounted cost Markov Decision Process (MDP). While MCTS is believed to provide an approximate value function for a given state with en…
Bayesian approach uses Gaussian process for reinforcement learning.
This paper offers a financial economic perspective on the optimal time (and age) at which the owner of a Variable Annuity (VA) policy with a Guaranteed Living Withdrawal Benefit (GLWB) rider should initiate guaranteed lifetime income payments. We abstract from utility, bequest and consumption preference issues by treat…
Curriculum learning speeds up agent learning in Minecraft, a complex visual domain.
Algorithm learns robust equilibrium in online Markov games with interactive data.
In recent years, deep reinforcement learning has been shown to be adept at solving sequential decision processes with high-dimensional state spaces such as in the Atari games. Many reinforcement learning problems, however, involve high-dimensional discrete action spaces as well as high-dimensional state spaces. This pa…
A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UCB exploration policy, and proved it has nearly optimal regret bound for finite-horizon episodic MDP. In this paper, we adapt Q-learning with…
The abstract reviews Markov models in life insurance surplus.
Although exploration in reinforcement learning is well understood from a theoretical point of view, provably correct methods remain impractical. In this paper we study the interplay between exploration and approximation, what we call approximate exploration. Our main goal is to further our theoretical understanding of …
This paper provides an empirical evaluation of recently developed exploration algorithms within the Arcade Learning Environment (ALE). We study the use of different reward bonuses that incentives exploration in reinforcement learning. We do so by fixing the learning algorithm used and focusing only on the impact of the…
We present in this paper a new premium computation principle based on the use of prior information from multiple sources for computing the premium charged to a policyholder. Under this framework, based on the use of Ordered Weighted Averaging (OWA) operators, we propose alternative collective and Bayes premiums and des…
The paper optimizes risk-sensitive RL with CVaR, achieving near-minimax-optimal results.
The paper analyzes GMWB annuities in low interest rate environments.
New algorithm achieves asymptotically optimal regret without horizon dependence.
Unified framework for distributional regret in bandits and reinforcement learning.
New algorithm achieves data-dependent regret bounds in MDPs with unknown transitions.
The "dancing metric" is a pseudo-riemannian metric of signature on the space of non-incident point-line pairs in the real projective plane . The null-curves of are given by the "dancing condition": the point is moving towards a point on the line, about which the li…
Study shows how to motivate AI agents like humans through learning from demonstrations.
We constructively prove the existence of time-discrete consumption processes for stochastic money accounts that fulfill a pre-specified positively homogeneous projection property (PHPP) and let the account always be positive and exactly zero at the end. One possible example is consumption rates forming a martingale und…
Improved online Q-learning for MDPs with concentration bounds.
New algorithm tackles non-stationary RL with near-optimal regret bounds.
We address the problem of correcting group discriminations within a score function, while minimizing the individual error. Each group is described by a probability density function on the set of profiles. We first solve the problem analytically in the case of two populations, with a uniform bonus-malus on the zones whe…
In distributional reinforcement learning (RL), the estimated distribution of value function models both the parametric and intrinsic uncertainties. We propose a novel and efficient exploration method for deep RL that has two components. The first is a decaying schedule to suppress the intrinsic uncertainty. The second …
In this paper we address the following question: Can we approximately sample from a Bayesian posterior distribution if we are only allowed to touch a small mini-batch of data-items for every sample we generate?. An algorithm based on the Langevin equation with stochastic gradients (SGLD) was previously proposed to solv…
Bayes-UCBVI tackles reinforcement learning with a new upper confidence bound method.
Many practical environments contain catastrophic states that an optimal agent would visit infrequently or never. Even on toy problems, Deep Reinforcement Learning (DRL) agents tend to periodically revisit these states upon forgetting their existence under a new policy. We introduce intrinsic fear (IF), a learned reward…
A framework disentangles controllable objects from visual signals for improved RL.
The explore{exploit dilemma is one of the central challenges in Reinforcement Learning (RL). Bayesian RL solves the dilemma by providing the agent with information in the form of a prior distribution over environments; however, full Bayesian planning is intractable. Planning with the mean MDP is a common myopic approxi…
This paper produces explicit strongly Hermitian Einstein-Maxwell solutions on the smooth compact -manifolds that are -bundles over compact Riemann surfaces of any genus. This generalizes the existence results by C. LeBrun in arXiv:1411.3992 and arXiv:1504.06669. Moreover, by calculating the (normalized) Einstei…
We define a manifold where objects are curves, which we parameterize as (, is the circle). Given a curve , we define the tangent space of at including in it all deformations of . In this paper we study geometries on the manifold of curves, pr…