This work analyzes how often to update the target network in Q-learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Multi-step methods such as Retrace() and -step -learning have become a crucial component of modern deep reinforcement learning agents. These methods are often evaluated as a part of bigger architectures and their evaluations rarely include enough samples to draw statistically significant conclusions about thei…
Spectral Adaptive Conformal Prediction for Structured Non-Exchangeable Data
In target tracking, the estimation of an unknown weaving target frequency is crucial for improving the miss distance. The estimation process is commonly carried out in a Kalman framework. The objective of this paper is to examine the potential of using neural networks in target tracking applications. To that end, we pr…
New measure of interference helps understand and mitigate learning issues in reinforcement learning.
In this paper, we address the fundamental problem of line spectral estimation in a Bayesian framework. We target model order and parameter estimation via variational inference in a probabilistic model in which the frequencies are continuous-valued, i.e., not restricted to a grid; and the coefficients are governed by a …
We develop a novel method for training of GANs for unsupervised and class conditional generation of images, called Linear Discriminant GAN (LD-GAN). The discriminator of an LD-GAN is trained to maximize the linear separability between distributions of hidden representations of generated and targeted samples, while the …
This paper analyzes how periodic and soft target updates stabilize linear Q-learning.
Newton's method converges faster than gradient descent in overparameterized neural networks.
FreSh shifts model's initial frequency spectrum to match target signal, improving neural representation performance.
In this paper, we propose a phase shift deep neural network (PhaseDNN) which provides a wideband convergence in approximating a high dimensional function during its training of the network. The PhaseDNN utilizes the fact that many DNN achieves convergence in the low frequency range first, thus, a series of moderately-s…
A new update rule for deep reinforcement learning reduces learning variance and variance in reference signals.
Sound event detection systems typically consist of two stages: extracting hand-crafted features from the raw audio waveform, and learning a mapping between these features and the target sound events using a classifier. Recently, the focus of sound event detection research has been mostly shifted to the latter stage usi…
Paper proposes a DRL-based controller for networked AP systems that reduces communication frequency.
Microstructure of market dynamics is studied through analysis of tick price data. Linear trend is introduced as a tool for such analysis. Trend arbitrage inequality is developed and tested. The inequality sets limiting relationship between trend, bid-ask spread, market reaction and average update frequency of price inf…
FFN addresses spectral bias in neural value approximation, improving reinforcement learning performance.
Motivated by recently published methods using frequency decompositions of convolutions (e.g. Octave Convolutions), we propose a novel convolution scheme to stabilize the training and reduce the likelihood of a mode collapse. The basic idea of our approach is to split convolutional filters into additive high and low fre…
New algorithm reduces FL sample and communication costs.
PQ-learning improves Q-learning by periodically updating target estimates.
For large scale on-line inference problems the update strategy is critical for performance. We derive an adaptive scan Gibbs sampler that optimizes the update frequency by selecting an optimum mini-batch size. We demonstrate performance of our adaptive batch-size Gibbs sampler by comparing it against the collapsed Gibb…
MPTE uses Transformer attention to estimate mixed-frequency factor models.
Radio frequency (RF) sensors are used alongside other sensing modalities to provide rich representations of the world. Given the high variability of complex-valued target responses, RF systems are susceptible to attacks masking true target characteristics from accurate identification. In this work, we evaluate differen…
One way to recognise an object is to study how the echo has been shaped during the interaction with the target. Wideband sonar allows the study of the energy distribution for a large range of frequencies. The frequency distribution contains information about an object, including its inner structure. This information is…
Paper extends SI method for detecting CPs in complex systems' frequency domain.
Paper improves policy updates in reinforcement learning to speed up learning.
We study the problem of optimizing the betting frequency in a dynamic game setting using Kelly's celebrated expected logarithmic growth criterion as the performance metric. The game is defined by a sequence of bets with independent and identically distributed returns X(k). The bettor selects the fraction of wealth K wa…
Permutation approach is suggested as a method to investigate financial time series in micro scales. The method is used to see how high frequency trading in recent years has affected the micro patterns which may be seen in financial time series. Tick to tick exchange rates are considered as examples. It is seen that var…
Improved HGF networks avoid negative precision errors in volatility updates.
This work analyzes how frequency components affect CNN predictions and robustness.
Large-scale machine learning training, in particular distributed stochastic gradient descent, needs to be robust to inherent system variability such as node straggling and random communication delays. This work considers a distributed training framework where each worker node is allowed to perform local model updates a…
Deep neural networks with convolutional layers usually process the entire spectrogram of an audio signal with the same time-frequency resolutions, number of filters, and dimensionality reduction scale. According to the constant-Q transform, good features can be extracted from audio signals if the low frequency bands ar…
This paper investigates Shampoo's heuristics and decouples preconditioner updates.
Three methods for tuning HMC diagonal scale matrices compared.
Deeper neural networks learn lower frequency functions faster, according to a new principle.
Federated learning systems are vulnerable to attacks from malicious clients. As the central server in the system cannot govern the behaviors of the clients, a rogue client may initiate an attack by sending malicious model updates to the server, so as to degrade the learning performance or enforce targeted model poisoni…
The paper proposes a method for valid multi-target regression predictions.
New method models complex dynamics using a base variable.
New model reduces volatility parameters and complexity.
In this paper, we propose a method for training neural networks when we have a large set of data with weak labels and a small amount of data with true labels. In our proposed model, we train two neural networks: a target network, the learner and a confidence network, the meta-learner. The target network is optimized to…
AI traders learn to exploit meta-orders from slower traders, increasing their profits.
We study the training process of Deep Neural Networks (DNNs) from the Fourier analysis perspective. We demonstrate a very universal Frequency Principle (F-Principle) -- DNNs often fit target functions from low to high frequencies -- on high-dimensional benchmark datasets such as MNIST/CIFAR10 and deep neural networks s…
It remains a puzzle that why deep neural networks (DNNs), with more parameters than samples, often generalize well. An attempt of understanding this puzzle is to discover implicit biases underlying the training process of DNNs, such as the Frequency Principle (F-Principle), i.e., DNNs often fit target functions from lo…
Muon outperforms GD in associative memory learning by balancing frequency components.
Paper tackles joint community detection and phase synchronization in stochastic block models.
EGMU optimizes portfolios using KL divergence, ensuring positive solutions.
Study efficient rebalancing strategies for portfolio tracking error.
GAIT-prop derives a biologically plausible learning rule from backpropagation.
This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask processing in the short-time Fourier transform (STFT)-domain is a typical speech enhancement method. To re…