A new method for optimizing deep neural networks using TKFAC.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New methods use Kronecker-factored approximations for faster deep learning optimization.
SINGD improves KFAC for memory-efficiency and stability in low-precision training.
Second-order optimization methods such as natural gradient descent have the potential to speed up training of neural networks by correcting for the curvature of the loss function. Unfortunately, the exact natural gradient is impractical to compute for large models, and most approximations either require an expensive it…
This paper studies iteration convergence of Kronecker graphical lasso (KGLasso) algorithms for estimating the covariance of an i.i.d. Gaussian random sample under a sparse Kronecker-product covariance model and MSE convergence rates. The KGlasso model, originally called the transposable regularized covariance model by …
A new optimization method reduces memory and compute requirements for deep learning.
A new iterative K-FAC algorithm reduces training time and memory usage.
A key challenge for gradient based optimization methods in model-free reinforcement learning is to develop an approach that is sample efficient and has low variance. In this work, we apply Kronecker-factored curvature estimation technique (KFAC) to a recently proposed gradient estimator for control variate optimization…
We introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks. The method is grounded in a Bayesian online learning framework, where we recursively approximate the posterior after every task with a Gaussian, leading to a quadratic penalty on changes to the we…
Develops efficient quasi-Newton methods for training deep neural networks.
This paper investigates Shampoo's heuristics and decouples preconditioner updates.
Optimization algorithms that leverage gradient covariance information, such as variants of natural gradient descent (Amari, 1998), offer the prospect of yielding more effective descent directions. For models with many parameters, the covariance matrix they are based on becomes gigantic, making them inapplicable in thei…
Improved continual learning for neural networks with BN layers using K-FAC extension.
A Kronecker product model is the set of visible marginal probability distributions of an exponential family whose sufficient statistics matrix factorizes as a Kronecker product of two matrices, one for the visible variables and one for the hidden variables. We estimate the dimension of these models by the maximum rank …
This paper speeds up K-FAC for deep learning by focusing on only a few eigen-modes.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
Reducing the test time resource requirements of a neural network while preserving test accuracy is crucial for running inference on resource-constrained devices. To achieve this goal, we introduce a novel network reparameterization based on the Kronecker-factored eigenbasis (KFE), and then apply Hessian-based structure…
K-FAC speeds up training of modern neural networks with linear weight-sharing.
Researchers solve the realization of Jordan-Kronecker invariants in Lie algebras.
Enhances Deep Hedging with K-FAC for financial data.
New methods improve Fisher Matrix approximations for neural networks at low cost.
Jointly models cause-of-death mortality rates across multiple countries and genders.
Gaussian Conditional Random Fields (GCRF), as a structured regression model, is designed to achieve higher regression accuracy than unstructured predictors at the expense of execution time, taking into account the objects similarities and the outputs of unstructured predictors simultaneously. As most structural models,…
A new method for learning Bayesian neural networks using layerwise inference.
New algorithm estimates matrix-valued regression parameters efficiently.
Despite all the impressive advances of recurrent neural networks, sequential data is still in need of better modelling. Truncated backpropagation through time (TBPTT), the learning algorithm most widely used in practice, suffers from the truncation bias, which drastically limits its ability to learn long-term dependenc…
Study on rotational hypersurfaces with constant Gauss-Kronecker curvature.
Machine learning predicts Kronecker coefficients with high accuracy.
Kronecker Products (KP) have been used to compress IoT RNN Applications by 15-38x compression factors, achieving better results than traditional compression methods. However when KP is applied to large Natural Language Processing tasks, it leads to significant accuracy loss (approx 26%). This paper proposes a way to re…
Bayesian method estimates Kronecker graphical models from autoregressive processes.
The wavelet Maximum Entropy on the Mean (wMEM) approach to the MEG inverse problem is revisited and extended to infer brain activity from full space-time data. The resulting dimensionality increase is tackled using a collection of techniques , that includes time and space dimension reduction (using respectively wavelet…
TensorSketch is an oblivious linear sketch introduced in Pagh'13 and later used in Pham, Pagh'13 in the context of SVMs for polynomial kernels. It was shown in Avron, Nguyen, Woodruff'14 that TensorSketch provides a subspace embedding, and therefore can be used for canonical correlation analysis, low rank approximation…
Method learns invariances in deep nets without human validation.
EiGLasso speeds up sparse Kronecker-sum covariance estimation.
Scalable Gaussian processes with latent Kronecker structure for large datasets.
In this paper we consider the use of the space vs. time Kronecker product decomposition in the estimation of covariance matrices for spatio-temporal data. This decomposition imposes lower dimensional structure on the estimated covariance matrix, thus reducing the number of samples required for estimation. To allow a sm…
We investigate 3-dimensional complete minimal hypersurfaces in the hyperbolic space with Gauss-Kronecker curvature identically zero. More precisely, we give a classification of complete minimal hypersurfaces with Gauss-Kronecker curvature identically zero, nowhere vanishing second fundamental form and …
The present paper discusses that a prescribed Gauss-Kronecker curvature problem on the product of unit spheres.
We investigate the structure of 3-dimensional complete minimal hypersurfaces in the unit sphere with Gauss-Kronecker curvature identically zero.
Efficiently models learning curves using Gaussian processes with latent Kronecker structure.
How can we model networks with a mathematically tractable model that allows for rigorous analysis of network properties? Networks exhibit a long list of surprising properties: heavy tails for the degree distribution; small diameters; and densification and shrinking diameters over time. Most present network models eithe…
We consider the problem of matrix approximation and denoising induced by the Kronecker product decomposition. Specifically, we propose to approximate a given matrix by the sum of a few Kronecker products of matrices, which we refer to as the Kronecker product approximation (KoPA). Because the Kronecker product is an ex…
The study predicts Kronecker coefficients using interpretable machine learning models.
We consider the problem of recovering a low-rank tensor from its noisy observation. Previous work has shown a recovery guarantee with signal to noise ratio for recovering a th order rank one tensor of size by recursive unfolding. In this paper, we first improve…
Recurrent Neural Networks (RNN) can be difficult to deploy on resource constrained devices due to their size.As a result, there is a need for compression techniques that can significantly compress RNNs without negatively impacting task accuracy. This paper introduces a method to compress RNNs for resource constrained e…
Dictionary learning and component analysis models are fundamental for learning compact representations that are relevant to a given task (feature extraction, dimensionality reduction, denoising, etc.). The model complexity is encoded by means of specific structure, such as sparsity, low-rankness, or nonnegativity. Unfo…
LQF linearizes deep models for better interpretability.
We investigate complete minimal hypersurfaces in the Euclidean space , with Gauss-Kronecker curvature identically zero. We prove that, if is a complete minimal hypersurface with Gauss-Kronecker curvature identically zero, nowhere vanishing second fundamental form and scalar curvature b…