Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

80160240320 · Jun 202019922001200920172026
48 results for bonus terms

We study an exploration method for model-free RL that generalizes the counter-based exploration bonus methods and takes into account long term exploratory value of actions rather than a single step look-ahead. We propose a model-free RL method that modifies Delayed Q-learning and utilizes the long-term exploration bonu…

2018-08-31abs ↗pdf ↗

Study minimax optimal RL in factored MDPs with bonus exploration.

problem Optimal reinforcement learning in episodic factored MDPs.
method Proposes two model-based algorithms with bonus exploration for minimax optimal regret.
result Achieves minimax optimal regret guarantees for rich factored structures.

The paper deals with bonus-malus systems with different claim types and varying deductibles. The premium relativities are softened for the policyholders who are in the malus zone and these policyholders are subject to per claim deductibles depending on their levels in the bonus-malus scale and the types of the reported…

2017-07-04abs ↗pdf ↗

New algorithm reduces reinforcement learning complexity, approaching contextual bandits.

problem Episodic reinforcement learning's difficulty compared to contextual bandits.
method Proposes MVP algorithm with a new Bernstein-type bonus for episodic reinforcement learning.
result Achieves near-optimal regret bound of $O\left(\left(\sqrt{SAK} + S^2A ight) \poly\log \left(SAHK ight) ight)$, improving state-of-the-art results.

Develops a Bonus-Malus model for cyber risk insurance to incentivize cybersecurity.

problem Lack of effective insurance strategies to incentivize cybersecurity.
method Proposes a Bonus-Malus model and a mathematical model with a numerical algorithm.
result Demonstrates how a Bonus-Malus system resolves moral hazard and benefits the insurer.

We discuss the pricing methodology for Bonus Certificates and Barrier Reverse-Convertible Structured Products. Pricing for a European barrier condition is straightforward for products of both types and depends on an efficient interpolation of observed market option pricing. Pricing products We discuss the pricing metho…

2016-07-31abs ↗pdf ↗

We introduce an exploration bonus for deep reinforcement learning methods that is easy to implement and adds minimal overhead to the computation performed. The bonus is the error of a neural network predicting features of the observations given by a fixed randomly initialized neural network. We also introduce a method …

2018-10-30abs ↗pdf ↗

Rewards are sparse in the real world and most of today's reinforcement learning algorithms struggle with such sparsity. One solution to this problem is to allow the agent to create rewards for itself - thus making rewards dense and more suitable for learning. In particular, inspired by curious behaviour in animals, obs…

2018-10-04abs ↗pdf ↗

The study analyzes how bonus-malus systems and delayed claims settlement affect insurance companies' financial stability.

problem Analyzing the impact of bonus-malus systems and delayed claims settlement on insurance companies' financial stability.
method Examined a discrete-time risk model with time-varying premiums, evaluating two types of claims and settlement delays.
result Delayed settlement of by-claims leads to lower ruin probabilities under specific assumptions.

The paper introduces Bellman-consistent pessimism to improve offline reinforcement learning without overly pessimistic bias.

problem Offline reinforcement learning's challenge of discovering good policies without exhaustive exploration.
method Introduces Bellman-consistent pessimism for function approximation, improving sample complexity and adaptability.
result Improves sample complexity by O(d)\mathcal{O}(d) in the action space finite case, and automatically adapts to bias-variance tradeoff.

Paper analyzes strategic underreporting in competitive insurance markets.

problem Strategic underreporting by insureds in competitive insurance markets.
method Develops a dynamic insurance market model with two competing companies and a continuum of insureds, examines the interaction between strategic underreporting and competitive pricing under a Bonus-Malus System framework.
result Establishes the existence and uniqueness of the insureds' optimal reporting barrier and its dependence on BMS premiums; proves the existence of Nash equilibrium premium strategies.

New UCB algorithm for learning PSRs with tractable computation and accuracy.

problem Learning predictive state representations in sequential decision-making problems.
method Proposes a novel UCB-type algorithm with a bonus term to estimate PSRs accurately and efficiently.
result First known UCB-type approach for PSRs with guaranteed model accuracy and computational tractability.

The study finds solar terms significantly impact China's stock market returns and volatility.

problem Investigating the effect of solar terms on China's stock market.
method Regression framework, analyzing multiple solar terms and their impact on return and volatility.
result Solar terms 1, 3, and 4 cause significant positive returns, while 8, 11, and 14 bring high volatility.

In this work, we consider the popular tree-based search strategy within the framework of reinforcement learning, the Monte Carlo Tree Search (MCTS), in the context of infinite-horizon discounted cost Markov Decision Process (MDP). While MCTS is believed to provide an approximate value function for a given state with en…

2019-02-14abs ↗pdf ↗

Curriculum learning speeds up agent learning in Minecraft, a complex visual domain.

problem Training agents to learn multiple tasks in a complex, visual domain.
method Learning-progress based curriculum and dynamic exploration bonuses.
result Curriculum learning improves agent performance in a complex reinforcement learning problem.

Algorithm learns robust equilibrium in online Markov games with interactive data.

problem Sim-to-real gap in reinforcement learning.
method Distributionally robust RL with minimum value assumption, least square value iteration.
result Sample-efficient algorithm for robust equilibrium in online Markov games.

Although exploration in reinforcement learning is well understood from a theoretical point of view, provably correct methods remain impractical. In this paper we study the interplay between exploration and approximation, what we call approximate exploration. Our main goal is to further our theoretical understanding of …

2018-08-29abs ↗pdf ↗

We present in this paper a new premium computation principle based on the use of prior information from multiple sources for computing the premium charged to a policyholder. Under this framework, based on the use of Ordered Weighted Averaging (OWA) operators, we propose alternative collective and Bayes premiums and des…

2015-11-12abs ↗pdf ↗

The paper optimizes risk-sensitive RL with CVaR, achieving near-minimax-optimal results.

problem Optimizing risk-sensitive reinforcement learning with CVaR objective.
method Developed algorithms for multi-arm bandits and online RL in MDPs, achieving near-minimax-optimal regret.
result Achieved near-minimax-optimal regret of O(τ1SAK)O(τ^{-1}\sqrt{SAK}) for constant ττ.

New algorithm achieves asymptotically optimal regret without horizon dependence.

problem Horizon-free regret minimization for reinforcement learning.
method Proposes a new algorithm and proves a regret upper bound.
result Regret upper bound of \(\tilde O(\sqrt{SAK} + S^8A^3)\) with failure probability \(\delta\).

Unified framework for distributional regret in bandits and reinforcement learning.

problem Characterizing the distribution of regret in multi-armed bandits and reinforcement learning.
method Unified framework with a UCBVI-style algorithm and distributional regret bounds.
result Distributional regret bounds with optimal trade-offs between expected and distributional regret.

New algorithm achieves data-dependent regret bounds in MDPs with unknown transitions.

problem Achieving best-of-both-worlds guarantees with data-dependent regret bounds in MDPs with unknown transitions.
method Optimistic follow-the-regularized-leader algorithm with new optimistic Q-function estimators and transition bonus.
result First-order, second-order, and path-length bounds with polylog(T) regret in the stochastic regime.

The "dancing metric" is a pseudo-riemannian metric g\pmb{g} of signature (2,2)(2,2) on the space M4M^4 of non-incident point-line pairs in the real projective plane RP2\mathbb{RP}^2. The null-curves of (M4,g)(M^4,\pmb{g}) are given by the "dancing condition": the point is moving towards a point on the line, about which the li…

2015-05-30abs ↗pdf ↗

We constructively prove the existence of time-discrete consumption processes for stochastic money accounts that fulfill a pre-specified positively homogeneous projection property (PHPP) and let the account always be positive and exactly zero at the end. One possible example is consumption rates forming a martingale und…

2007-11-27abs ↗pdf ↗

Improved online Q-learning for MDPs with concentration bounds.

problem Online Q-learning in infinite-horizon discounted MDPs with sublinear regret for large gaps.
method Smoothed εnε_n-Greedy exploration scheme combining εnε_n-greedy and Boltzmann exploration, analyzed using concentration bounds for contractive Markovian stochastic approximation.
result Near-ildeO(N9/10) ilde{O}(N^{9/10}) regret bound for Smoothed εnε_n-Greedy exploration scheme.

New algorithm tackles non-stationary RL with near-optimal regret bounds.

problem Model-free reinforcement learning in non-stationary Markov decision processes.
method Proposed RestartQ-UCB algorithm with Freedman-type bonus terms.
result Achieves near-optimal dynamic regret bound in non-stationary RL.

We address the problem of correcting group discriminations within a score function, while minimizing the individual error. Each group is described by a probability density function on the set of profiles. We first solve the problem analytically in the case of two populations, with a uniform bonus-malus on the zones whe…

2018-06-07abs ↗pdf ↗

In distributional reinforcement learning (RL), the estimated distribution of value function models both the parametric and intrinsic uncertainties. We propose a novel and efficient exploration method for deep RL that has two components. The first is a decaying schedule to suppress the intrinsic uncertainty. The second …

2019-05-13abs ↗pdf ↗

In this paper we address the following question: Can we approximately sample from a Bayesian posterior distribution if we are only allowed to touch a small mini-batch of data-items for every sample we generate?. An algorithm based on the Langevin equation with stochastic gradients (SGLD) was previously proposed to solv…

2012-06-27abs ↗pdf ↗

Bayes-UCBVI tackles reinforcement learning with a new upper confidence bound method.

problem Optimizing exploration in reinforcement learning without bonuses.
method Bayes-UCBVI uses a quantile of a Q-value function posterior as an upper confidence bound.
result Proves a regret bound of order O~(H3SAT)\widetilde{O}(\sqrt{H^3SAT}) for tabular reinforcement learning.

Many practical environments contain catastrophic states that an optimal agent would visit infrequently or never. Even on toy problems, Deep Reinforcement Learning (DRL) agents tend to periodically revisit these states upon forgetting their existence under a new policy. We introduce intrinsic fear (IF), a learned reward…

2016-11-03abs ↗pdf ↗

A framework disentangles controllable objects from visual signals for improved RL.

problem Improving sample efficiency and game performance in vision-based RL.
method Action-conditioned video prediction to disentangle controllable objects.
result Improved sample efficiency and game performance in Atari games.

The explore{exploit dilemma is one of the central challenges in Reinforcement Learning (RL). Bayesian RL solves the dilemma by providing the agent with information in the form of a prior distribution over environments; however, full Bayesian planning is intractable. Planning with the mean MDP is a common myopic approxi…

2012-03-15abs ↗pdf ↗

This paper produces explicit strongly Hermitian Einstein-Maxwell solutions on the smooth compact 44-manifolds that are S2S^2-bundles over compact Riemann surfaces of any genus. This generalizes the existence results by C. LeBrun in arXiv:1411.3992 and arXiv:1504.06669. Moreover, by calculating the (normalized) Einstei…

2015-11-21abs ↗pdf ↗

We define a manifold MM where objects cMc\in M are curves, which we parameterize as c:S1Rnc:S^1\to R^n (n2n\ge 2, S1S^1 is the circle). Given a curve cc, we define the tangent space TcMT_cM of MM at cc including in it all deformations h:S1Rnh:S^1\to R^n of cc. In this paper we study geometries on the manifold of curves, pr…

2006-04-30abs ↗pdf ↗