Paper optimizes energy-based controller for swinging up a pendulum using entropy search.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper solves pendulum swing-up problem using RL.
We study valuation of swing options on commodity markets when the commodity prices are driven by multiple factors. The factors are modeled as diffusion processes driven by a multidimensional Lévy process. We set up a valuation model in terms of a dynamic programming problem where the option can be exercised continuousl…
This paper compares linear regression and neural networks for pricing swing options.
Paper finds significant impact of stock market swings on equity risk premium predictability.
The paper solves complex swing option pricing equations with numerical methods.
Paper introduces a new volatility model for natural gas markets and discusses swing option pricing.
New method uses neural networks for optimal stopping time problems.
Two methods for pricing swing contracts using neural networks or explicit functions.
In this paper, we investigate a numerical algorithm for the pricing of swing options, relying on the so-called optimal quantization method. The numerical procedure is described in details and numerous simulations are provided to assert its efficiency. In particular, we carry out a comparison with the Longstaff-Schwartz…
Study on convex ordering in stochastic control for swing contracts, proving value function convexity.
We study an optimal control problem related to swing option pricing in a general non-Markovian setting in continuous time. As a main result we show that the value process solves a first-order non-linear backward stochastic partial differential equation. Based on this result we can characterize the set of optimal contro…
The paper models natural gas futures prices and volatility, using Monte Carlo and reinforcement learning.
The paper introduces and studies hedging for game (Israeli) style extension of swing options considered as multiple exercise derivatives. Assuming that the underlying security can be traded without restrictions we derive a formula for valuation of multiple exercise options via classical hedging arguments. Introducing t…
This paper provides fast estimates for complex option types.
We characterize the value of swing contracts in continuous time as the unique viscosity solution of a Hamilton-Jacobi-Bellman equation with suitable boundary conditions. The case of contracts with penalties is straightforward, and in that case only a terminal condition is needed. Conversely, the case of contracts with …
We introduce a new probabilistic method for solving a class of impulse control problems based on their representations as Backward Stochastic Differential Equations (BSDEs for short) with constrained jumps. As an example, our method is used for pricing Swing options. We deal with the jump constraint by a penalization p…
In Bender and Dokuchaev (2013), we studied a control problem related to swing option pricing in a general non-Markovian setting. The main result there shows that the value process of this control problem can be uniquely characterized in terms of a first order backward SPDE and a pathwise differential inclusion. In the …
Designing optimal controllers continues to be challenging as systems are becoming complex and are inherently nonlinear. The principal advantage of reinforcement learning (RL) is its ability to learn from the interaction with the environment and provide optimal control strategy. In this paper, RL is explored in the cont…
We present a data-efficient reinforcement learning algorithm resistant to observation noise. Our method extends the highly data-efficient PILCO algorithm (Deisenroth & Rasmussen, 2011) into partially observed Markov decision processes (POMDPs) by considering the filtering process during policy evaluation. PILCO conduct…
Swing options on the gas market are american style option where daily quantities exercices are constrained and global quantities exerciced each year constrained too. The option holder has to decide each day how much he consumes of the quantities satisfying the constraints and tries to use a strategy in order to maximiz…
Controller seeks informative system observations to predict nonlinear dynamics.
A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go. These games have a discrete action space. However, many real-world reinforcement learning domains have continuous action spaces, for example in robotic control, na…
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algo…
In this paper we study perpetual American call and put options in an exponential Lévy model. We consider a negative effective discount rate which arises in a number of financial applications including stock loans and real options, where the strike price can potentially grow at a higher rate than the original discount f…
We give several new positive finite presentations for the pure braid group that are easy to remember and simple in form. All of our presentations involve a metric on the punctured disc so that the punctures are arranged "convexly", which is why we describe them as geometric presentaitons. Motivated by a presentation fo…
We present an approach to identify concise equations from data using a shallow neural network approach. In contrast to ordinary black-box regression, this approach allows understanding functional relations and generalizing them from observed data to unseen parts of the parameter space. We show how to extend the class o…
Proposes a transfer learning framework to improve U.S. election prediction models.
Paper presents fast methods for pricing energy derivatives using mean-reverting jump-diffusion models.
The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. Among the few proposed approaches, the recent…
mm-Pose detects human skeletons in real-time using mmWave radar and CNNs.
Empowerment quantifies the influence an agent has on its environment. This is formally achieved by the maximum of the expected KL-divergence between the distribution of the successor state conditioned on a specific action and a distribution where the actions are marginalised out. This is a natural candidate for an intr…
Deep Q-Learning models optimal exercise strategies for option-type products.
Model learns and plans in real-time under constraints for robotic systems.
Physics-informed learning framework for pH systems and EB-PBC control.
We start briefly surveying research on optimal stopping games since their introduction by E.B.Dynkin more than 40 years ago. Recent renewed interest to dynkin's games is due, in particular, to the study of Israeli (game) options introduced in 2000. We discuss the work on these options and related derivative securities …
Bayesian optimization outperforms other methods in hyperparameter tuning for reinforcement learning.
Reinforcement Learning methods are capable of solving complex problems, but resulting policies might perform poorly in environments that are even slightly different. In robotics especially, training and deployment conditions often vary and data collection is expensive, making retraining undesirable. Simulation training…
We use probabilistic methods to characterise time dependent optimal stopping boundaries in a problem of multiple optimal stopping on a finite time horizon. Motivated by financial applications we consider a payoff of immediate stopping of "put" type and the underlying dynamics follows a geometric Brownian motion. The op…
Symbolic regression constructs simple equations for complex systems.
Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.
Proposes DLGPD model to learn dynamics from images for planning.
Symbolic regression constructs smooth value functions for reinforcement learning.
CryptoGAT improves cryptocurrency price prediction by treating it as a graph problem.
PhI-GPR improves power grid state estimation and forecasting.
We establish several new stylised facts concerning the intra-day seasonalities of stock dynamics. Beyond the well known U-shaped pattern of the volatility, we find that the average correlation between stocks increases throughout the day, leading to a smaller relative dispersion between stocks. Somewhat paradoxically, t…
New neural network approximates convex option prices.
In this comment we discuss the problem of reconciling the linear efficiency of price returns with the long-memory of supply and demand. We present new evidence that shows that efficiency is maintained by a liquidity imbalance that co-moves with the imbalance of buyer vs. seller initiated transactions. For example, duri…