A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
This paper presents several models addressing optimal portfolio choice, optimal portfolio liquidation, and optimal portfolio transition issues, in which the expected returns of risky assets are unknown. Our approach is based on a coupling between Bayesian learning and dynamic programming techniques that leads to partia…
Robots can rapidly acquire new skills from demonstrations. However, during generalisation of skills or transitioning across fundamentally different skills, it is unclear whether the robot has the necessary knowledge to perform the task. Failing to detect missing information often leads to abrupt movements or to collisi…
A critical and challenging problem in reinforcement learning is how to learn the state-action value function from the experience replay buffer and simultaneously keep sample efficiency and faster convergence to a high quality solution. In prior works, transitions are uniformly sampled at random from the replay buffer o…
As part of Basel II's incremental risk charge (IRC) methodology, this paper summarizes our extensive investigations of constructing transition probability matrices (TPMs) for unsecuritized credit products in the trading book. The objective is to create monthly or quarterly TPMs with predefined sectors and ratings that …
Novel framework for risk-sensitive reinforcement learning with robustness against uncertainty.
problem Risk-sensitive reinforcement learning with uncertainty in transition dynamics.
method Developed a risk-sensitive robust Markov decision process (RSRMDP), derived its Bellman equation, and proposed a Bayesian Dynamic Programming (Bayesian DP) algorithm.
result Demonstrated convergence to near-optimal policies and analyzed sample and computational complexities.
A Robust Markov Decision Process (RMDP) is a sequential decision making model that accounts for uncertainty in the parameters of dynamic systems. This uncertainty introduces difficulties in learning an optimal policy, especially for environments with large state spaces. We propose two algorithms, RTD-DQN and Deep-RoK, …
Robust Markov Decision Processes (RMDPs) intend to ensure robustness with respect to changing or adversarial system behavior. In this framework, transitions are modeled as arbitrary elements of a known and properly structured uncertainty set and a robust optimal policy can be derived under the worst-case scenario. In t…
We here present a model of the dynamics of extremism based on opinion dynamics in order to understand the circumstances which favour its emergence and development in large fractions of the general public. Our model is based on the bounded confidence hypothesis and on the evolution of initially anti-conformist agents to…
The goal of system identification is to learn about underlying physics dynamics behind the time-series data. To model the probabilistic and nonparametric dynamics model, Gaussian process (GP) have been widely used; GP can estimate the uncertainty of prediction and avoid over-fitting. Traditional GPSSMs, however, are ba…
We generalize the recently proposed quantum model for the stock market by Zhang and Huang to make it consistent with the discrete nature of the stock price. In this formalism, the price of the stock and its trend satisfy the generalized uncertainty relation and the corresponding generalized Hamiltonian contains an addi…
Minimum energy paths for transitions such as atomic and/or spin rearrangements in thermalized systems are the transition paths of largest statistical weight. Such paths are frequently calculated using the nudged elastic band method, where an initial path is iteratively shifted to the nearest minimum energy path. The co…
Robust reinforcement learning aims to produce policies that have strong guarantees even in the face of environments/transition models whose parameters have strong uncertainty. Existing work uses value-based methods and the usual primitive action setting. In this paper, we propose robust methods for learning temporally …
We present an analysis of oil prices in US$ and in other major currencies that diagnoses unsustainable faster-than-exponential behavior. This supports the hypothesis that the recent oil price run-up has been amplified by speculative behavior of the type found during a bubble-like expansion. We also attempt to unravel t…