Advantage amplification helps RL in slow-evolving latent-state environments.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the non-stationary stochastic multiarmed bandit (MAB) problem and propose two generic algorithms, namely, the limited memory deterministic sequencing of exploration and exploitation (LM-DSEE) and the Sliding-Window Upper Confidence Bound# (SW-UCB#). We rigorously analyze these algorithms in abruptly-changing a…
In this paper, we study the portfolio optimization problem with general utility functions and when the return and volatility of underlying asset are slowly varying. An asymptotic optimal strategy is provided within a specific class of admissible controls under this problem setup. Specifically, we first establish a rigo…
Paper proposes faster adaptation to distribution shifts in online settings.
Deep reinforcement learning methods attain super-human performance in a wide range of environments. Such methods are grossly inefficient, often taking orders of magnitudes more data than humans to achieve reasonable performance. We propose Neural Episodic Control: a deep reinforcement learning agent that is able to rap…
Market inefficiencies arise from density-dependent returns in a noisy environment.
FADE adapts machine learning models to evolving data efficiently.
We examine how the structure of the world trade network has been shaped by globalization and recessions over the last 40 years. We show that by treating the world trade network as an evolving system, theory predicts the trade network is more sensitive to evolutionary shocks and recovers more slowly from them now than i…
This paper studies the portfolio optimization problem when the investor's utility is general and the return and volatility of the risky asset are fast mean-reverting, which are important to capture the fast-time scale in the modeling of stock price volatility. Motivated by the heuristic derivation in [J.-P. Fouque, R. …
Two adaptive algorithms improve tracking regret in dynamic expert advice problems.
Rough stochastic volatility models have attracted a lot of attentions recently, in particular for the linear option pricing problem. In this paper, starting with power utilities, we propose to use a martingale distortion representation of the optimal value function for the nonlinear asset allocation problem in a (non-M…
We show that the emergence of systemic risk in complex systems can be understood from the evolution of functional networks representing interactions inferred from fluctuation correlations between macroscopic observables. Specifically, we analyze the long-term collective dynamics of the New York Stock Exchange between 1…
New method detects anomalies in computing centers' logs.
We describe an optimal adversarial attack formulation against autoregressive time series forecast using Linear Quadratic Regulator (LQR). In this threat model, the environment evolves according to a dynamical system; an autoregressive model observes the current environment state and predicts its future values; an attac…
A new algorithm for non-stationary linear bandits with improved regret bound.
In classical contagion models, default systems are Markovian conditionally on the observation of their stochastic environment, with interacting intensities. This necessitates that the environment evolves autonomously and is not influenced by the history of the default events. We extend the classical literature and allo…
Study non-stationary MDPs using worst-case RL, proposing RATS algorithm.
A new neural model evolves to learn at the synaptic level.
MPC outperforms reactive budgeting in non-stationary return environments.
Study on slow convergence in geometric variational problems.
Study shows subordinated Cramér-Lundberg model increases ruin probability.
Bayesian method for dynamic correlation matrices improves accuracy and responsiveness.
LC-SAC tackles non-stationary dynamics in reinforcement learning.
A new method simulates large, diverse populations of learning agents evolving in games.
Unique global solutions found for specific initial data.
An adaptive clustering algorithm learns from evolving data without manual tuning.
This paper describes a reference architecture for self-maintaining systems that can learn continually, as data arrives. In environments where data evolves, we need architectures that manage Machine Learning (ML) models in production, adapt to shifting data distributions, cope with outliers, retrain when necessary, and …
A method to optimize deep networks by sequentially minimizing risk functions.
ParsNet tackles weakly supervised data streams with a self-evolving deep neural network.
Fractional stochastic volatility models have been widely used to capture the non-Markovian structure revealed from financial time series of realized volatility. On the other hand, empirical studies have identified scales in stock price volatility: both fast-time scale on the order of days and slow-scale on the order of…
Bayesian Federated Learning improves model reliability in dynamic environments.
A method to improve time series forecasting by dynamically adjusting weights of forecasters.
Partially performative prediction studies how predictive models influence future data.
Trading strategies evolve in a simulated market to outperform real data.
A new approach for deep exploration in sparse reward reinforcement learning.
Solving statistical learning problems often involves nonconvex optimization. Despite the empirical success of nonconvex statistical optimization methods, their global dynamics, especially convergence to the desirable local minima, remain less well understood in theory. In this paper, we propose a new analytic paradigm …
This work develops agents to learn generalizable policies for dynamic network environments.
FSNet improves online time series forecasting by balancing fast adaptation and old knowledge.
The paper proposes a method to adapt machine learning models to changing conditions.
Image recognition using Deep Learning has been evolved for decades though advances in the field through different settings is still a challenge. In this paper, we present our findings in searching for better image classifiers in offline and online environments. We resort to Convolutional Neural Network and its variatio…
The paper explains how continuous language models can produce discrete, interpretable meanings.
We introduce the active exploration problem in Markov decision processes (MDPs). Each state of the MDP is characterized by a random value and the learner should gather samples to estimate the mean value of each state as accurately as possible. Similarly to active exploration in multi-armed bandit (MAB), states may have…
New algorithm for non-stationary bandits with slow drifts.
The generative learning phase of Autoencoder (AE) and its successor Denosing Autoencoder (DAE) enhances the flexibility of data stream method in exploiting unlabelled samples. Nonetheless, the feasibility of DAE for data stream analytic deserves in-depth study because it characterizes a fixed network capacity which can…
New method learns optimal environment and goal difficulty for reinforcement learning.
Let M be a closed orientable 3-manifold with a negatively curved Riemannian metric. Let {M_i} be a collection of finite regular covers with degree d_i. (1) If the Heegaard genus of M_i grows more slowly than the square root of d_i, then M_i has positive first Betti number for all sufficiently large i. (2) The strong He…
Empirical studies indicate the presence of multi-scales in the volatility of underlying assets: a fast-scale on the order of days and a slow-scale on the order of months. In our previous works, we have studied the portfolio optimization problem in a Markovian setting under each single scale, the slow one in [Fouque and…
Click-through rate~(CTR) prediction, whose goal is to estimate the probability of the user clicks, has become one of the core tasks in advertising systems. For CTR prediction model, it is necessary to capture the latent user interest behind the user behavior data. Besides, considering the changing of the external envir…