This paper investigates the effectiveness of decoupled weight decay at the start of training.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Decoupled GCN is shown to be equivalent to label propagation.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
We identity a by-far-unrecognized problem of Adam-style optimizers which results from unnecessary coupling between momentum and adaptivity. The coupling leads to instability and divergence when the momentum and adaptivity parameters are mismatched. In this work, we propose a method, Laprop, which decouples momentum and…
The paper decouples shrinkage and selection in Bayesian Quantile Regression.
From a simplified analysis of adaptive methods, we derive AvaGrad, a new optimizer which outperforms SGD on vision tasks when its adaptability is properly tuned. We observe that the power of our method is partially explained by a decoupling of learning rate and adaptability, greatly simplifying hyperparameter search. I…
Unsupervised representation learning via generative modeling is a staple to many computer vision applications in the absence of labeled data. Variational Autoencoders (VAEs) are powerful generative models that learn representations useful for data generation. However, due to inherent challenges in the training objectiv…
We document a mechanism operating in complex adaptive systems leading to dynamical pockets of predictability (``prediction days''), in which agents collectively take predetermined courses of action, transiently decoupled from past history. We demonstrate and test it out-of-sample on synthetic minority and majority game…
We develop a novel family of algorithms for the online learning setting with regret against any data sequence bounded by the empirical Rademacher complexity of that sequence. To develop a general theory of when this type of adaptive regret bound is achievable we establish a connection to the theory of decoupling inequa…
Inner product-based convolution has been a central component of convolutional neural networks (CNNs) and the key to learning visual representations. Inspired by the observation that CNN-learned features are naturally decoupled with the norm of features corresponding to the intra-class variation and the angle correspond…
A new method routes EEG covariance matrices across domains using adaptive subspace selection.
We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, machine translation, and language modeling, it performs on par or better than well tuned SGD with mom…
Study separates learning rate effects from adaptive gradient methods.
Regularized training of an autoencoder typically results in hidden unit biases that take on large negative values. We show that negative biases are a natural result of using a hidden layer whose responsibility is to both represent the input data and act as a selection mechanism that ensures sparsity of the representati…
This paper addresses the problem of learning the optimal control policy for a nonlinear stochastic dynamical system with continuous state space, continuous action space and unknown dynamics. This class of problems are typically addressed in stochastic adaptive control and reinforcement learning literature using model-b…
This paper investigates Shampoo's heuristics and decouples preconditioner updates.
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
Spectral decoupling improves neural network generalization in medical imaging.
We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients. While neural network optimization has previously been studied using natural gradient methods which are closely related to Bayesian inference, they were unable to recover standard optim…
Gradient-based meta-learning techniques are both widely applicable and proficient at solving challenging few-shot learning and fast adaptation problems. However, they have practical difficulties when operating on high-dimensional parameter spaces in extreme low-data regimes. We show that it is possible to bypass these …
Nash integrates covariate-specific side info into sparse regression via neural networks.
We investigate probabilistic decoupling of labels supplied for training, from the underlying classes for prediction. Decoupling enables an inference scheme general enough to implement many classification problems, including supervised, semi-supervised, positive-unlabelled, noisy-label and suggests a general solution to…
The paper proposes a decoupled approach to efficiently estimate CoVaR, a measure of systemic financial risk.
Clarifies method of phase synchronization for decoupling linear differential equations.
Kahler geometry explains decoupling of Kerr perturbations.
SA-BCP combines long-term and local evidence for efficient, adaptive online prediction.
Lo-Hp decouples weight generation into local and global policies to improve flexibility and efficiency.
Recent variants improve knowledge distillation performance.
New algorithms for interpreting complex multivariate functions.
New method combines spectral and sparse methods for Gaussian processes.
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
It is well known that in stochastic multi-armed bandits (MAB), the sample mean of an arm is typically not an unbiased estimator of its true mean. In this paper, we decouple three different sources of this selection bias: adaptive \emph{sampling} of arms, adaptive \emph{stopping} of the experiment, and adaptively \emph{…
Unified framework suppresses model bias in semi-supervised learning with decoupled sampling control.
Gradient-EM Bayesian meta-learning accelerates adaptation with reduced computation and improved robustness.
Current reinforcement learning (RL) methods can successfully learn single tasks but often generalize poorly to modest perturbations in task domain or training procedure. In this work, we present a decoupled learning strategy for RL that creates a shared representation space where knowledge can be robustly transferred. …
This paper studies an entropy-based multi-objective Bayesian optimization (MBO). The entropy search is successful approach to Bayesian optimization. However, for MBO, existing entropy-based methods ignore trade-off among objectives or introduce unreliable approximations. We propose a novel entropy-based MBO called Pare…
Stylized facts can be regarded as constraints for any modeling attempt of price dynamics on a financial market, in that an empirically reasonable model has to reproduce these stylized facts at least qualitatively. The dynamics of market prices is modeled on a macro-level as the result of the dynamic coupling of two dyn…
Backpropagation algorithm is indispensable for the training of feedforward neural networks. It requires propagating error gradients sequentially from the output layer all the way back to the input layer. The backward locking in backpropagation algorithm constrains us from updating network layers in parallel and fully l…
In this paper we target the problem of transferring policies across multiple environments with different dynamics parameters and motor noise variations, by introducing a framework that decouples the processes of policy learning and system identification. Efficiently transferring learned policies to an unknown environme…
A new model decouples global and local image representations without supervision.
New equations simplify gauge-theoretic Khovanov homology solutions.
A new method for anomaly detection adapts to local non-stationarity in low-data regimes.
Decoupled PFNs improve sequential decision-making by separating epistemic and aleatoric uncertainties.
Novel Adam-family method with decoupled weight decay for training neural networks.
This paper studies the convergence behaviour of dictionary learning via the Iterative Thresholding and K-residual Means (ITKrM) algorithm. On one hand it is proved that ITKrM is a contraction under much more relaxed conditions than previously necessary. On the other hand it is shown that there seem to exist stable fixe…
DAPS++ improves diffusion-based image restoration by decoupling prior and likelihood.
In this paper we propose to solve an important problem in recommendation -- user cold start, based on meta leaning method. Previous meta learning approaches finetune all parameters for each new user, which is both computing and storage expensive. In contrast, we divide model parameters into fixed and adaptive parts and…
A new framework reduces inconsistencies in chaotic surrogate modeling.