OMGD algorithm optimizes online convex optimization with switching costs and delayed gradients.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper improves competitive and dynamic regret bounds for smoothed online learning.
SCaLE tackles dynamic regret in noisy bandit feedback with switching costs.
New algorithm reduces switching costs in multinomial logit bandit problems.
Combines dynamic programming and neural networks for optimal portfolio execution in regime-switching markets.
In this paper, we consider a discrete time economy where we assume that the short term interest rate follows a quadratic term structure of a regime switching asset process. The possible non-linear structure and the fact that the interest rate can have different economic or financial trends justify the interest of Regim…
The paper explores dynamic regret with switching cost in online decision making.
The problem of optimal switching between nonlinear autonomous subsystems is investigated in this study where the objective is not only bringing the states to close to the desired point, but also adjusting the switching pattern, in the sense of penalizing switching occurrences and assigning different preferences to util…
New RL algorithm reduces policy switching cost to loglog(T) with similar regret.
Paper presents an efficient algorithm for linear MDP with low switching cost.
LaMBO optimizes modular systems with switching costs, achieving better results than existing methods.
This paper tackles near-optimal adversarial RL with switching costs, providing algorithms and matching lower bounds.
Adaptive Bayesian Optimization for resource-constrained experiments with switching costs.
Algorithm for bandits with switching costs achieves optimal regret bounds.
We study online learning when partial feedback information is provided following every action of the learning process, and the learner incurs switching costs for changing his actions. In this setting, the feedback information system can be represented by a graph, and previous works studied the expected regret of the le…
Study optimal liquidation strategies with infinite horizon and regime switching.
We consider the problem of the optimal trading strategy in the presence of linear costs, and with a strict cap on the allowed position in the market. Using Bellman's backward recursion method, we show that the optimal strategy is to switch between the maximum allowed long position and the maximum allowed short position…
New algorithm reduces switching costs in RL beyond linear MDPs.
We consider the exploration-exploitation tradeoff in linear quadratic (LQ) control problems, where the state dynamics is linear and the cost function is quadratic in states and controls. We analyze the regret of Thompson sampling (TS) (a.k.a. posterior-sampling for reinforcement learning) in the frequentist setting, i.…
New RL algorithms reduce costs for single-agent and federated learning.
Study tackles balancing policy switching costs in offline RL.
We study the power of different types of adaptive (nonoblivious) adversaries in the setting of prediction with expert advice, under both full-information and bandit feedback. We measure the player's performance using a new notion of regret, also known as policy regret, which better captures the adversary's adaptiveness…
We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization. This is motivated by the difficulty of running fully adaptive algorithms in real-world applications (such as medical domains), and …
We introduce and solve a new type of quadratic backward stochastic differential equation systems defined in an infinite time horizon, called \emph{ergodic BSDE systems}. Such systems arise naturally as candidate solutions to characterize forward performance processes and their associated optimal trading strategies in a…
New algorithm reduces RL complexity with low switching costs.
The present work studies and analyzes general defaultable OTC contract in presence of a contingent CSA, which is a theoretical counterparty risk mitigation mechanism of switching type that allows the counterparty of a general OTC contract to switch from zero to full/perfect collateralization and switch back whenever sh…
Paper solves complex game theory problems with new equations.
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
This paper analyzes the problem of starting and stopping a Cox-Ingersoll-Ross (CIR) process with fixed costs. In addition, we also study a related optimal switching problem that involves an infinite sequence of starts and stops. We establish the conditions under which the starting-stopping and switching problems admit …
Optimizes dividend payouts with fixed costs and regime switching.
We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given switching budget. For this problem, we prove matching upper and lower bounds on the optimal (i.e., minimax) regret, and provide efficient rate-o…
This paper is concerned with cost optimization of an insurance company. The surplus of the insurance company is modeled by a controlled regime switching diffusion, where the regime switching mechanism provides the fluctuations of the random environment. The goal is to find an optimal control that minimizes the total co…
Study optimal portfolios in a non-Markovian regime-switching model with random time horizon.
New algorithms tackle adversarial combinatorial bandits with switching costs.
Study when to replace machine learning models with new data.
The paper solves a complex control problem with stochastic elements and switching conditions.
The study examines portfolio optimization with quadratic transaction costs, complicating the optimization process.
Paper tackles online optimization with memory and competitive control.
New algorithm reduces decision switching in dynamic environments.
We solve non-Markovian optimal switching problems in discrete time on an infinite horizon, when the decision maker is risk aware and the filtration is general, and establish existence and uniqueness of solutions for the associated reflected backward stochastic difference equations. An example application to hydropower …
This paper uses recent results on continuous-time finite-horizon optimal switching problems with negative switching costs to prove the existence of a saddle point in an optimal stopping (Dynkin) game. Sufficient conditions for the game's value to be continuous with respect to the time horizon are obtained using recent …
We study the adversarial multi-armed bandit problem where partial observations are available and where, in addition to the loss incurred for each action, a \emph{switching cost} is incurred for shifting to a new action. All previously known results incur a factor proportional to the independence number of the feedback …
We study the solution's existence for a generalized Dynkin game of switching type which is shown to be the natural representation for general defaultable OTC contract with contingent CSA. This is a theoretical counterparty risk mitigation mechanism that allows the counterparty of a general OTC contract to switch from z…
This paper studies the optimal VIX futures trading problems under a regime-switching model. We consider the VIX as mean reversion dynamics with dependence on the regime that switches among a finite number of states. For the trading strategies, we analyze the timings and sequences of the investor's market participation,…
The paper solves a utility-based hedging problem with quadratic costs.
Study how transaction costs impact stock returns and holdings in equilibrium.
We reconsider the problem of optimal trading in the presence of linear and quadratic costs, for arbitrary linear costs but in the limit where quadratic costs are small. Using matched asymptotic expansion techniques, we find that the trading speed vanishes inside a band that is narrower than in the absence of quadratic …
In this paper, we consider a risk-based optimal investment problem of an insurer in a regime-switching jump diffusion model with noisy memory. Using the model uncertainty modeling, we formulate the investment problem as a zero-sum, stochastic differential delay game between the insurer and the market, with a convex ris…