New algorithm reduces switching costs in multinomial logit bandit problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The problem of optimal switching between nonlinear autonomous subsystems is investigated in this study where the objective is not only bringing the states to close to the desired point, but also adjusting the switching pattern, in the sense of penalizing switching occurrences and assigning different preferences to util…
New RL algorithm reduces policy switching cost to loglog(T) with similar regret.
Paper presents an efficient algorithm for linear MDP with low switching cost.
LaMBO optimizes modular systems with switching costs, achieving better results than existing methods.
As a metric to measure the performance of an online method, dynamic regret with switching cost has drawn much attention for online decision making problems. Although the sublinear regret has been provided in many previous researches, we still have little knowledge about the relation between the dynamic regret and the s…
This paper tackles near-optimal adversarial RL with switching costs, providing algorithms and matching lower bounds.
Adaptive Bayesian Optimization for resource-constrained experiments with switching costs.
Algorithm for bandits with switching costs achieves optimal regret bounds.
We study online learning when partial feedback information is provided following every action of the learning process, and the learner incurs switching costs for changing his actions. In this setting, the feedback information system can be represented by a graph, and previous works studied the expected regret of the le…
SCaLE tackles dynamic regret in noisy bandit feedback with switching costs.
OMGD algorithm optimizes online convex optimization with switching costs and delayed gradients.
The paper improves competitive and dynamic regret bounds for smoothed online learning.
New algorithm reduces switching costs in RL beyond linear MDPs.
New RL algorithms reduce costs for single-agent and federated learning.
Study tackles balancing policy switching costs in offline RL.
We study the power of different types of adaptive (nonoblivious) adversaries in the setting of prediction with expert advice, under both full-information and bandit feedback. We measure the player's performance using a new notion of regret, also known as policy regret, which better captures the adversary's adaptiveness…
We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization. This is motivated by the difficulty of running fully adaptive algorithms in real-world applications (such as medical domains), and …
New algorithm reduces RL complexity with low switching costs.
The present work studies and analyzes general defaultable OTC contract in presence of a contingent CSA, which is a theoretical counterparty risk mitigation mechanism of switching type that allows the counterparty of a general OTC contract to switch from zero to full/perfect collateralization and switch back whenever sh…
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
This paper analyzes the problem of starting and stopping a Cox-Ingersoll-Ross (CIR) process with fixed costs. In addition, we also study a related optimal switching problem that involves an infinite sequence of starts and stops. We establish the conditions under which the starting-stopping and switching problems admit …
Optimizes dividend payouts with fixed costs and regime switching.
We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given switching budget. For this problem, we prove matching upper and lower bounds on the optimal (i.e., minimax) regret, and provide efficient rate-o…
This paper is concerned with cost optimization of an insurance company. The surplus of the insurance company is modeled by a controlled regime switching diffusion, where the regime switching mechanism provides the fluctuations of the random environment. The goal is to find an optimal control that minimizes the total co…
New algorithms tackle adversarial combinatorial bandits with switching costs.
Study when to replace machine learning models with new data.
Combines dynamic programming and neural networks for optimal portfolio execution in regime-switching markets.
New algorithm reduces decision switching in dynamic environments.
We solve non-Markovian optimal switching problems in discrete time on an infinite horizon, when the decision maker is risk aware and the filtration is general, and establish existence and uniqueness of solutions for the associated reflected backward stochastic difference equations. An example application to hydropower …
This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …
We study the adversarial multi-armed bandit problem where partial observations are available and where, in addition to the loss incurred for each action, a \emph{switching cost} is incurred for shifting to a new action. All previously known results incur a factor proportional to the independence number of the feedback …
We study the solution's existence for a generalized Dynkin game of switching type which is shown to be the natural representation for general defaultable OTC contract with contingent CSA. This is a theoretical counterparty risk mitigation mechanism that allows the counterparty of a general OTC contract to switch from z…
This paper studies the optimal VIX futures trading problems under a regime-switching model. We consider the VIX as mean reversion dynamics with dependence on the regime that switches among a finite number of states. For the trading strategies, we analyze the timings and sequences of the investor's market participation,…
In this work we study the price-hedge issue for general defaultable contracts characterized by the presence of a contingent CSA of switching type. This is a contingent risk mitigation mechanism that allow the counterparties of a defaultable contract to switch from zero to full/perfect collateralization and switch back …
Efficiently optimize GPs by reusing candidate solutions multiple times.
The rBergomi model is improved with a regime switching change of measure to match market VIX smiles.
This paper studies the timing of trades under mean-reverting price dynamics subject to fixed transaction costs. We solve an optimal double stopping problem to determine the optimal times to enter and subsequently exit the market, when prices are driven by an exponential Ornstein-Uhlenbeck process. In addition, we analy…
This paper studies a finite-fuel two-dimensional degenerate singular stochastic control problem under regime switching that is motivated by the optimal irreversible extraction problem of an exhaustible commodity. A company extracts a natural resource from a reserve with finite capacity, and sells it in the market at a …
The paper explores how sinks and diagonal patterns prevent attention oversmoothing.
Algorithm selects best model based on state, reducing costs.
Breaks down complex nonlinear dynamics into simpler components.
This paper studies four trading algorithms of a professional trader at a multilateral trading facility, observing a realistic two-sided limit order book whose dynamics are driven by the order book events. The identity of the trader can be either internalizing or regular, either a hedge fund or a brokery agency. The spe…
Here we shall consider a very popular practical applied problem of managing mode switching (in this work we are considering managing billing plans). Out of the two parties (service provider and service consumer), participating in the processes modelled here, we shall consider only a consumer type of a problem. Herein w…
This paper improves Q-learning bounds using reference-advantage decomposition.
Investigates JM for reducing downside risk in market regimes.
A network of agents attempt to learn some unknown state of the world drawn by nature from a finite set. Agents observe private signals conditioned on the true state, and form beliefs about the unknown state accordingly. Each agent may face an identification problem in the sense that she cannot distinguish the truth in …
Combines MCTS and neural networks for efficient multi-period financial planning.