Develops a learning model predictive controller for competitive racing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper analyzes and proposes a new stopping criterion for recursive Bayesian classification.
This paper establishes the existence of a unique nonnegative continuous viscosity solution to the HJB equation associated with a Markovian linear-quadratic control problems with singular terminal state constraint and possibly unbounded cost coefficients. The existence result is based on a novel comparison principle for…
Survive method improves model-based RL by avoiding terminal states, reducing sample complexity.
New reward function improves GAIL performance in task-based environments.
We establish existence, uniqueness and regularity of solution results for a class of backward stochastic partial differential equations with singular terminal condition. The equation describes the value function of non-Markovian stochastic optimal control problem in which the terminal state of the controlled process is…
Paper argues the bear case for Bitcoin is bounded and terminal states are neutral to positive.
Humans tend to learn complex abstract concepts faster if examples are presented in a structured manner. For instance, when learning how to play a board game, usually one of the first concepts learned is how the game ends, i.e. the actions that lead to a terminal state (win, lose or draw). The advantage of learning end-…
Deep reinforcement learning has achieved great successes in recent years, but there are still open challenges, such as convergence to locally optimal policies and sample inefficiency. In this paper, we contribute a novel self-supervised auxiliary task, i.e., Terminal Prediction (TP), estimating temporal closeness to te…
We analyze linear McKean-Vlasov forward-backward SDEs arising in leader-follower games with mean-field type control and terminal state constraints on the state process. We establish an existence and uniqueness of solutions result for such systems in time-weighted spaces as well as a {convergence} result of the solution…
Resource allocation improved using machine learning from terminal positions.
Solves steering problem with continuous time, Hilbert-Schmidt cost, and matrix ODEs.
TVM improves generative modeling by matching terminal velocities.
In this paper, the `Approximate Message Passing' (AMP) algorithm, initially developed for compressed sensing of signals under i.i.d. Gaussian measurement matrices, has been extended to a multi-terminal setting (MAMP algorithm). It has been shown that similar to its single terminal counterpart, the behavior of MAMP algo…
Investors optimize their portfolios within a Wasserstein ball to match a benchmark's risk profile.
We provide a probabilistic solution of a not necessarily Markovian control problem with a state constraint by means of a Backward Stochastic Differential Equation (BSDE). The novelty of our solution approach is that the BSDE possesses a singular terminal condition. We prove that a solution of the BSDE exists, thus part…
In the continuous time mean-variance model, we want to minimize the variance (risk) of the investment portfolio with a given mean at terminal time. However, the investor can stop the investment plan at any time before the terminal time. To solve this kind of problem, we consider to minimize the variances of the investm…
Optimal asset allocation strategy outperforms stochastic benchmark.
Counterfactual Regret Minimization (CFR) has found success in settings like poker which have both terminal states and perfect recall. We seek to understand how to relax these requirements. As a first step, we introduce a simple algorithm, local no-regret learning (LONR), which uses a Q-learning-like update rule to allo…
In many environments only a tiny subset of all states yield high reward. In these cases, few of the interactions with the environment provide a relevant learning signal. Hence, we may want to preferentially train on those high-reward states and the probable trajectories leading to them. To this end, we advocate for the…
New method learns diffusion bridges for rare events.
A new framework models multi-state events and biomarkers.
To improve the efficient frontier of the classical mean-variance model in continuous time, we propose a varying terminal time mean-variance model with a constraint on the mean value of the portfolio asset, which moves with the varying terminal time. Using the embedding technique from stochastic optimal control in conti…
We consider the class of short rate interest rate models for which the short rate is proportional to the exponential of a Gaussian Markov process x(t) in the terminal measure r(t) = a(t) exp(x(t)). These models include the Black, Derman, Toy and Black, Karasinski models in the terminal measure. We show that such intere…
In this work, we consider the problem of autonomously discovering behavioral abstractions, or options, for reinforcement learning agents. We propose an algorithm that focuses on the termination condition, as opposed to -- as is common -- the policy. The termination condition is usually trained to optimize a control obj…
Assuming that agents' preferences satisfy first-order stochastic dominance, we show how the Expected Utility paradigm can rationalize all optimal investment choices: the optimal investment strategy in any behavioral law-invariant (state-independent) setting corresponds to the optimum for an expected utility maximizer w…
We consider risk-averse agents who compete for liquidity in an Almgren--Chriss market impact model. Mathematically, this situation can be described by a Nash equilibrium for a certain linear-quadratic differential game with state constraints. The state constraints enter the problem as terminal boundary conditions f…
A framework solves parametric families of MFGs efficiently.
In this paper, we prove some convergence results of a special case of optimistic policy iteration algorithm for stochastic shortest path problem. We consider both Monte Carlo and methods for the policy evaluation step under the condition that the termination state will eventually be reached almost surely.
Adaptive stopping in MCMC using classifier-based dynamics
Study bounds for prices of European and American options with optional termination.
The paper finds optimal threshold strategies for insurance companies with a positive terminal value at creeping ruin.
New method preserves distances in time series data.
Proves finite step termination of Kähler-Einstein metric singularity formation.
We study a robust maximization problem from terminal wealth and consumption under a convex constraints on the portfolio. We state the existence and the uniqueness of the consumption-investment strategy by studying the associated quadratic backward stochastic differential equation (BSDE in short). We characterize the op…
The study proves a key inequality for specific types of three-dimensional spaces.
We present an extension of Monte Carlo Tree Search (MCTS) that strongly increases its efficiency for trees with asymmetry and/or loops. Asymmetric termination of search trees introduces a type of uncertainty for which the standard upper confidence bound (UCB) formula does not account. Our first algorithm (MCTS-T), whic…
Optimizes portfolio growth rate for a behavioral investor considering terminal relative growth rate.
New test for SGD in binary classification reduces computation time.
The aim of this contribution is to derive a general matrix formula for the net period premium paid in more than one state. For this purpose we propose to combine actuarial technics with the graph optimization methodology. The obtained result is useful for example to more advanced models of dread disease insurances allo…
Employee stock options (ESOs) are American-style call options that can be terminated early due to employment shock. This paper studies an ESO valuation framework that accounts for job termination risk and jumps in the company stock price. Under general Lévy stock price dynamics, we show that a higher job termination ri…
We prove that the sum of the -invariants of two different Kollár components of a Kawamata log terminal singularity is less than .
New method for computing terminal embeddings in sublinear time.
Locally adaptive clustering for tree delineation.
Is an option to early terminate a swap at its market value worth zero? At first sight it is, but in presence of counterparty risk it depends on the criteria used to determine such market value. In case of a single uncollateralised swap transaction under ISDA between two defaultable counterparties, the additional unilat…
This paper investigates optimal trading strategies in a financial market with multidimensional stock returns where the drift is an unobservable multivariate Ornstein-Uhlenbeck process. Information about the drift is obtained by observing stock returns and expert opinions. The latter provide unbiased estimates on the cu…
In reinforcement learning, a decision needs to be made at some point as to whether it is worthwhile to carry on with the learning process or to terminate it. In many such situations, stochastic elements are often present which govern the occurrence of rewards, with the sequential occurrences of positive rewards randoml…
Study optimal liquidation with multiple regimes using BSDEs with singular terminal values.