This paper analyzes how periodic and soft target updates stabilize linear Q-learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PQ-learning improves Q-learning by periodically updating target estimates.
Recently, the technique of local updates is a powerful tool in centralized settings to improve communication efficiency via periodical communication. For decentralized settings, it is still unclear how to efficiently combine local updates and decentralized communication. In this work, we propose an algorithm named as L…
This work analyzes how often to update the target network in Q-learning.
We introduce a new class of forward performance processes that are endogenous and predictable with regards to an underlying market information set and, furthermore, are updated at discrete times. We analyze in detail a binomial model whose parameters are random and updated dynamically as the market evolves. We show tha…
Federated learning is a distributed framework according to which a model is trained over a set of devices, while keeping data localized. This framework faces several systems-oriented challenges which include (i) communication bottleneck since a large number of devices upload their local updates to a parameter server, a…
ADER addresses continual learning in session-based recommendation by periodically replaying exemplars with adaptive distillation.
Step-DAD improves BED by periodically updating a design policy during experiments.
For a long investment time horizon, it is preferable to rebalance the portfolio weights at intermediate times. This necessitates a multi-period market model in which portfolio optimization is usually done through dynamic programming. However, this assumes a known distribution for the parameters of the financial time se…
We study collaborative machine learning (ML) across wireless devices, each with its own local dataset. Offloading these datasets to a cloud or an edge server to implement powerful ML solutions is often not feasible due to latency, bandwidth and privacy constraints. Instead, we consider federated edge learning (FEEL), w…
GD converges in unstable regimes, even with oscillatory behavior.
Local adaptive methods in FL can accelerate convergence but introduce bias, which is corrected.
Since August 2000, the stock market in the USA as well as most other western markets have depreciated almost in synchrony according to complex patterns of drops and local rebounds. In \cite{SZ02QF}, we have proposed to describe this phenomenon using the concept of a log-periodic power law (LPPL) antibubble, characteriz…
In this paper, we explore techniques centered around periodic sampling of model weights that provide convergence improvements on gradient update methods (vanilla \acs{SGD}, Momentum, Adam) for a variety of vision problems (classification, detection, segmentation). Importantly, our algorithms provide better, faster and …
Paper addresses OPE for dependent bandit samples using MDS and batch updates.
Holdout set improves risk score accuracy without biasing predictions.
Paper proposes a DRL-based controller for networked AP systems that reduces communication frequency.
Through the analysis of a dataset of ultra high frequency order book updates, we introduce a model which accommodates the empirical properties of the full order book together with the stylized facts of lower frequency financial data. To do so, we split the time interval of interest into periods in which a well chosen r…
Communication-efficient SGD algorithms, which allow nodes to perform local updates and periodically synchronize local models, are highly effective in improving the speed and scalability of distributed SGD. However, a rigorous convergence analysis and comparative study of different communication-reduction strategies rem…
Generalising the idea of the classical EM algorithm that is widely used for computing maximum likelihood estimates, we propose an EM-Control (EM-C) algorithm for solving multi-period finite time horizon stochastic control problems. The new algorithm sequentially updates the control policies in each time period using Mo…
Applicability of the concept of financial log-periodicity is discussed and encouragingly verified for various phases of the world stock markets development in the period 2000-2010. In particular, a speculative forecasting scenario designed in the end of 2004, that properly predicted the world stock market increases in …
Bayesian model predicts online activity participation.
Communication overhead is one of the key challenges that hinders the scalability of distributed optimization algorithms. In this paper, we study local distributed SGD, where data is partitioned among computation nodes, and the computation nodes perform local updates with periodically exchanging the model among the work…
Distributed optimization is essential for training large models on large datasets. Multiple approaches have been proposed to reduce the communication overhead in distributed training, such as synchronizing only after performing multiple local SGD steps, and decentralized methods (e.g., using gossip algorithms) to decou…
Federated CTMC model estimates bridge deterioration hazards without sharing raw data.
A framework for navigating environments with spatially correlated obstacles and uncertain blockage status.
Lambda Learner improves model freshness in data streams.
The use of target networks has been a popular and key component of recent deep Q-learning algorithms for reinforcement learning, yet little is known from the theory side. In this work, we introduce a new family of target-based temporal difference (TD) learning algorithms and provide theoretical analysis on their conver…
A new method for efficiently updating large-scale matrices in real-time.
Proposes a deep learning method for modeling dynamic individual-level latent trajectories with changing parameters.
This work introduces a fixed-point optimization for variational inference.
PER-ETD improves ETD by reducing variance to polynomial complexity.
Paper proposes an efficient method for calibrating spatio-temporal forecasts.
Large-scale machine learning training, in particular distributed stochastic gradient descent, needs to be robust to inherent system variability such as node straggling and random communication delays. This work considers a distributed training framework where each worker node is allowed to perform local model updates a…
The study analyzes sharpness dynamics in neural networks, revealing mechanisms and conditions.
Paper develops MMOT framework for financial applications with neural acceleration.
New PFPPs based on rank-dependent utility for better performance control.
This work explores Target Networks and Functional Regularization in deep Reinforcement Learning.
Despite the robust structure of the Internet, it is still susceptible to disruptive routing updates that prevent network traffic from reaching its destination. Our research shows that BGP announcements that are associated with disruptive updates tend to occur in groups of relatively high frequency, followed by periods …
Deep reinforcement learning (DRL) methods such as the Deep Q-Network (DQN) have achieved state-of-the-art results in a variety of challenging, high-dimensional domains. This success is mainly attributed to the power of deep neural networks to learn rich domain representations for approximating the value function or pol…
In this paper, we study the design and analysis of experiments conducted on a set of units over multiple time periods where the starting time of the treatment may vary by unit. The design problem involves selecting an initial treatment time for each unit in order to most precisely estimate both the instantaneous and cu…
Improved NiNo networks accelerate Adam training by up to 50%.
RSI uses Bayesian inference to monitor compliance in rule-governed domains.
The paper solves portfolio selection using Rényi divergence and optimization.
We address the problem of predicting spatio-temporal processes with temporal patterns that vary across spatial regions, when data is obtained as a stream. That is, when the training dataset is augmented sequentially. Specifically, we develop a localized spatio-temporal covariance model of the process that can capture s…
RIFLE improves deep transfer learning by reinitializing fully-connected layers.
A distributed algorithm for online multi-task learning reduces communication and runtime costs.
Inverse classification, the process of making meaningful perturbations to a test point such that it is more likely to have a desired classification, has previously been addressed using data from a single static point in time. Such an approach yields inflated probability estimates, stemming from an implicitly made assum…