Novel Adam-family method with decoupled weight decay for training neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Decoupled GCN is shown to be equivalent to label propagation.
Inner product-based convolution has been a central component of convolutional neural networks (CNNs) and the key to learning visual representations. Inspired by the observation that CNN-learned features are naturally decoupled with the norm of features corresponding to the intra-class variation and the angle correspond…
Deep learning has gained great popularity due to its widespread success on many inference problems. We consider the application of deep learning to the sparse linear inverse problem encountered in compressive sensing, where one seeks to recover a sparse signal from a small number of noisy linear measurements. In this p…
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
We develop a novel family of algorithms for the online learning setting with regret against any data sequence bounded by the empirical Rademacher complexity of that sequence. To develop a general theory of when this type of adaptive regret bound is achievable we establish a connection to the theory of decoupling inequa…
Spectral decoupling improves neural network generalization in medical imaging.
We propose a novel neural architecture search algorithm via reinforcement learning by decoupling structure and operation search processes. Our approach samples candidate models from the multinomial distribution on the policy vectors defined on the two search spaces independently. The proposed technique improves the eff…
This paper addresses the problem of learning the optimal control policy for a nonlinear stochastic dynamical system with continuous state space, continuous action space and unknown dynamics. This class of problems are typically addressed in stochastic adaptive control and reinforcement learning literature using model-b…
Proves energy estimates for tensorial wave equations, decoupling components for stability proof.
We propose a novel algorithm to solve the expectation propagation relaxation of Bayesian inference for continuous-variable graphical models. In contrast to most previous algorithms, our method is provably convergent. By marrying convergent EP ideas from (Opper&Winther 05) with covariance decoupling techniques (Wipf&Nag…
We investigate probabilistic decoupling of labels supplied for training, from the underlying classes for prediction. Decoupling enables an inference scheme general enough to implement many classification problems, including supervised, semi-supervised, positive-unlabelled, noisy-label and suggests a general solution to…
While Bayesian neural networks have many appealing characteristics, current priors do not easily allow users to specify basic properties such as expected lengthscale or amplitude variance. In this work, we introduce Poisson Process Radial Basis Function Networks, a novel prior that is able to encode amplitude stationar…
The paper proposes a decoupled approach to efficiently estimate CoVaR, a measure of systemic financial risk.
In this paper, we study the generalization properties of online learning based stochastic methods for supervised learning problems where the loss function is dependent on more than one training sample (e.g., metric learning, ranking). We present a generic decoupling technique that enables us to provide Rademacher compl…
Asynchronous parallel optimization algorithms for solving large-scale machine learning problems have drawn significant attention from academia to industry recently. This paper proposes a novel algorithm, decoupled asynchronous proximal stochastic gradient descent (DAP-SGD), to minimize an objective function that is the…
This paper investigates the effectiveness of decoupled weight decay at the start of training.
Clarifies method of phase synchronization for decoupling linear differential equations.
Large-scale Gaussian process inference has long faced practical challenges due to time and space complexity that is superlinear in dataset size. While sparse variational Gaussian process models are capable of learning from large-scale data, standard strategies for sparsifying the model can prevent the approximation of …
Kahler geometry explains decoupling of Kerr perturbations.
Generative model solves financial market equilibria with stable reinforcement learning.
AdaDEM decouples EM into two parts to improve class overlap and uncertainty.
Recent variants improve knowledge distillation performance.
New algorithms for interpreting complex multivariate functions.
New method combines spectral and sparse methods for Gaussian processes.
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
Decouples critic chunk length from policy to improve policy reactivity and performance.
Current reinforcement learning (RL) methods can successfully learn single tasks but often generalize poorly to modest perturbations in task domain or training procedure. In this work, we present a decoupled learning strategy for RL that creates a shared representation space where knowledge can be robustly transferred. …
Reinforcement learning frameworks have introduced abstractions to implement and execute algorithms at scale. They assume standardized simulator interfaces but are not concerned with identifying suitable task representations. We present Wield, a first-of-its kind system to facilitate task design for practical reinforcem…
Study proves global existence and decay for complex wave equations.
Stylized facts can be regarded as constraints for any modeling attempt of price dynamics on a financial market, in that an empirically reasonable model has to reproduce these stylized facts at least qualitatively. The dynamics of market prices is modeled on a macro-level as the result of the dynamic coupling of two dyn…
We study a generalized setup for learning from demonstration to build an agent that can manipulate novel objects in unseen scenarios by looking at only a single video of human demonstration from a third-person perspective. To accomplish this goal, our agent should not only learn to understand the intent of the demonstr…
Backpropagation algorithm is indispensable for the training of feedforward neural networks. It requires propagating error gradients sequentially from the output layer all the way back to the input layer. The backward locking in backpropagation algorithm constrains us from updating network layers in parallel and fully l…
A new model decouples global and local image representations without supervision.
New equations simplify gauge-theoretic Khovanov homology solutions.
Decoupled PFNs improve sequential decision-making by separating epistemic and aleatoric uncertainties.
DAPS++ improves diffusion-based image restoration by decoupling prior and likelihood.
SGD without replacement decouples into curvature-following and flatness-regularizing steps.
New method uses diffusion models for Bayesian inverse problems.
We investigate the convergence and stability properties of the decoupled extended Kalman filter learning algorithm (DEKF) within the long-short term memory network (LSTM) based online learning framework. For this purpose, we model DEKF as a perturbed extended Kalman filter and derive sufficient conditions for its stabi…
DAPS++ improves diffusion-based image restoration by decoupling prior and likelihood.
Lion optimizer performs well in training AI models with memory efficiency.
CAKD framework optimizes knowledge transfer by focusing on influential components of distillation.
FTPL policy achieves best-of-both-worlds regret in decoupled bandits with reduced computational cost.
New method detects heuristics in complex game strategies.
Deep fictitious play converges to Nash equilibrium in stochastic differential games.
This paper studies an entropy-based multi-objective Bayesian optimization (MBO). The entropy search is successful approach to Bayesian optimization. However, for MBO, existing entropy-based methods ignore trade-off among objectives or introduce unreliable approximations. We propose a novel entropy-based MBO called Pare…
Study compares constrained and decoupled moduli spaces of manifolds with particles and discs.