ADVISOR dynamically balances imitation and reinforcement learning to overcome the imitation gap.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm reduces regret for many bandit algorithms with logarithmic dependence on number of algorithms.
Reinforcement learning is a promising approach to synthesizing policies for challenging robotics tasks. A key problem is how to ensure safety of the learned policy---e.g., that a walking robot does not fall over or that an autonomous car does not run into an obstacle. We focus on the setting where the dynamics are know…
New algorithm reduces switching costs in multinomial logit bandit problems.
New polynomial invariants derived from birack and switch structures.
In this paper, we study optimal switching problems under ambiguity. To characterize the optimal switching under ambiguity in the finite horizon, we use multidimensional reflected backward stochastic differential equations (multidimensional RBSDEs) and show that a value function of the optimal switching under ambiguity …
This work extends identifiability analysis to sequential latent variable models, focusing on Switching Dynamical Systems.
The problem of optimal switching between nonlinear autonomous subsystems is investigated in this study where the objective is not only bringing the states to close to the desired point, but also adjusting the switching pattern, in the sense of penalizing switching occurrences and assigning different preferences to util…
Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched Speech and Natural Language Processing. We motivate why processing code-switched …
New algorithms improve sampling from complex distributions.
Squirrel switches between optimizers for better performance.
Study approximates financial market with discrete-time models.
This paper studies deep learning methodologies for portfolio optimization in the US equities market. We present a novel residual switching network that can automatically sense changes in market regimes and switch between momentum and reversal predictors accordingly. The residual switching network architecture combines …
Optimizes control of hybrid systems with multiple switching processes.
Study tackles balancing policy switching costs in offline RL.
New algorithm learns switching dynamics from multiple neural signals.
Paper tackles utility maximization with job-switching and retirement constraints.
New RL algorithm reduces policy switching cost to loglog(T) with similar regret.
In the recent years, the desire and need to understand sequential data has been increasing, with particular interest in sequential contexts such as patient monitoring, understanding daily activities, video surveillance, stock market and the like. Along with the constant flow of data, it is critical to classify and segm…
Proposes on-the-fly joint feature selection and classification for time-sensitive decisions.
Paper presents an efficient algorithm for linear MDP with low switching cost.
This paper studies the impact of limited switches on resource-constrained dynamic pricing with demand learning. We focus on the classical price-based blind network revenue management problem and extend our results to the bandits with knapsacks problem. In both settings, a decision maker faces stochastic and distributio…
Study strategic competition in commodity markets using impulse-switching controls.
One type of switch simplifies operations on lattice knots.
Optimal switching regret for all segmentations in online convex optimisation.
As a metric to measure the performance of an online method, dynamic regret with switching cost has drawn much attention for online decision making problems. Although the sublinear regret has been provided in many previous researches, we still have little knowledge about the relation between the dynamic regret and the s…
This paper tackles near-optimal adversarial RL with switching costs, providing algorithms and matching lower bounds.
This paper addresses parameter estimation for wave equations with Markovian switching.
Algorithm for bandits with switching costs achieves optimal regret bounds.
In this paper, we derive the family switching formula of -n two-sphere fiber bundle embedded in a smooth four-manifold fiber bundle. In the smooth category, it is a partial generalization of Fintushel-Stern's argument for four-manifolds. We also derive an algebraic analogue of the family switching formula, allowing the…
Regime switching volatility models provide a tractable method of modelling stochastic volatility. Currently the most popular method of regime switching calibration is the Hamilton filter. We propose using the Baum-Welch algorithm, an established technique from Engineering, to calibrate regime switching models instead. …
Audit fees change based on company and economic factors during auditor switching.
We study the problem of switching-constrained online convex optimization (OCO), where the player has a limited number of opportunities to change her action. While the discrete analog of this online learning task has been studied extensively, previous work in the continuous setting has neither established the minimax ra…
LaMBO optimizes modular systems with switching costs, achieving better results than existing methods.
A new framework tunes hyperparameters in real-time for contextual bandits.
New algorithm reduces switching costs in RL beyond linear MDPs.
Develops identifiability theory for multi-lag regime-switching models.
We study online learning when partial feedback information is provided following every action of the learning process, and the learner incurs switching costs for changing his actions. In this setting, the feedback information system can be represented by a graph, and previous works studied the expected regret of the le…
Paper derives analytical formulas for NLD-CEV moments with regime switching.
OMGD algorithm optimizes online convex optimization with switching costs and delayed gradients.
Pricing financial or real options with arbitrary payoffs in regime-switching models is an important problem in finance. Mathematically, it is to solve, under certain standard assumptions, a general form of optimal stopping problems in regime-switching models. In this article, we reduce an optimal stopping problem with …
This paper presents a novel technique based on gradient boosting to train the final layers of a neural network (NN). Gradient boosting is an additive expansion algorithm in which a series of models are trained sequentially to approximate a given function. A neural network can also be seen as an additive expansion where…
In this paper, we consider the problem of pricing discretely-sampled variance swaps based on a hybrid model of stochastic volatility and stochastic interest rate with regime-switching. Our modelling framework extends the Heston stochastic volatility model by including the CIR stochastic interest rate and model paramete…
New algorithm reduces RL complexity with low switching costs.
The paper deals with regression problems, in which the nonsmooth target is assumed to switch between different operating modes. Specifically, piecewise smooth (PWS) regression considers target functions switching deterministically via a partition of the input space, while switching regression considers arbitrary switch…
Many complex dynamical phenomena can be effectively modeled by a system that switches among a set of conditionally linear dynamical modes. We consider two such models: the switching linear dynamical system (SLDS) and the switching vector autoregressive (VAR) process. Our Bayesian nonparametric approach utilizes a hiera…
Label switching is a phenomenon arising in mixture model posterior inference that prevents one from meaningfully assessing posterior statistics using standard Monte Carlo procedures. This issue arises due to invariance of the posterior under actions of a group; for example, permuting the ordering of mixture components …
This paper first describes a class of uncertain stochastic control systems with Markovian switching, and derives an Itô-Liu formula for Markov-modulated processes. And we characterize an optimal control law, which satisfies the generalized Hamilton-Jacobi-Bellman (HJB) equation with Markovian switching. Then, by using …