Three methods for tuning HMC diagonal scale matrices compared.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
Diagonal transformations preserve independence structures in non-Gaussian distributions.
This paper optimizes diagonal preconditioning to improve matrix condition numbers.
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
The approximate joint diagonalization of a set of matrices consists in finding a basis in which these matrices are as diagonal as possible. This problem naturally appears in several statistical learning tasks such as blind signal separation. We consider the diagonalization criterion studied in a seminal paper by Pham (…
Develops a novel stochastic algorithm for diagonal estimation of large matrices.
Recurrent Neural Networks (RNNs) are designed to handle sequential data but suffer from vanishing or exploding gradients. Recent work on Unitary Recurrent Neural Networks (uRNNs) have been used to address this issue and in some cases, exceed the capabilities of Long Short-Term Memory networks (LSTMs). We propose a simp…
Method estimates M-matrices in graphical models with improved accuracy.
In this paper, we propose a new Recurrent Neural Network (RNN) architecture. The novelty is simple: We use diagonal recurrent matrices instead of full. This results in better test likelihood and faster convergence compared to regular full RNNs in most of our experiments. We show the benefits of using diagonal recurrent…
An algorithm for computing positive semidefinite factorizations of matrices.
This paper solves matrix blind joint block diagonalization with noise.
Despite their successes, what makes kernel methods difficult to use in many large scale problems is the fact that storing and computing the decision function is typically expensive, especially at prediction time. In this paper, we overcome this difficulty by proposing Fastfood, an approximation that accelerates such co…
A fast metric learning framework using Gershgorin disc alignment.
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
In this paper, we study deep diagonal circulant neural networks, that is deep neural networks in which weight matrices are the product of diagonal and circulant ones. Besides making a theoretical analysis of their expressivity, we introduced principled techniques for training these models: we devise an initialization s…
Gradient descent optimally trains RNNs without overparameterization.
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank represent…
Meta learns low-rank covariance factors for better uncertainty estimation.
The crossing matrix of a braid on strands is the integer matrix with zero diagonal whose entry is the algebraic number (positive minus negative) of crossings by strand over strand . When restricted to the subgroup of pure braids, this defines a homomorphism onto the additive subgroup of $N…
Diagonal linear networks converge to lasso regularization path during training.
New method tackles high-dimensional SBL without covariance matrices.
Localized sketching improves matrix multiplication and ridge regression complexity.
Non-orthogonal joint diagonalization (NJD) free of prewhitening has been widely studied in the context of blind source separation (BSS) and array signal processing, etc. However, NJD is used to retrieve the jointly diagonalizable structure for a single set of target matrices which are mostly formulized with a single da…
The paper explores spinors and polyforms using quaternions and octonions.
The paper proves a distribution claim for neural network Jacobians.
We study algebraic properties of matrices whose rows are mutual neighbours, and are also neigbours of 0 ("neighbour" in the sense of a certain nilpotency condition). The intended application is in synthetic differential geometry. For a square matrix of this kind, the product of the diagonal entries equals the determina…
This work is motivated by numerical solutions to Hamilton-Jacobi-Bellman quasi-variational inequalities (HJBQVIs) associated with combined stochastic and impulse control problems. In particular, we consider (i) direct control, (ii) penalized, and (iii) semi-Lagrangian discretization schemes applied to the HJBQVI proble…
Study on likelihood functions, associative equations, and Frobenius manifolds.
Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…
Researchers develop geodesics for a new metric on correlation matrices.
This note classifies splittable lattices in a specific Lie group.
We present explicit formulas for the coordinates in which the Hamiltonians of the Benenti systems with flat metrics take natural form and the metrics in question are represented by constant diagonal matrices.
A new metric learning framework for signed graphs using Gershgorin disc alignment.
Chevalley theorems extended to isotropic functions on matrix spaces.
In this paper we consider the use of the space vs. time Kronecker product decomposition in the estimation of covariance matrices for spatio-temporal data. This decomposition imposes lower dimensional structure on the estimated covariance matrix, thus reducing the number of samples required for estimation. To allow a sm…
Improves BBVI for high-dimensional Gaussian approximations by using low-rank approximations.
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
New method estimates sparse covariance matrices in logit mixtures.
D-LinOSS models learn to dissipate energy, improving performance on long-range tasks.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
The paper models financial correlation matrices using permutation invariant Gaussian models and predicts market anomalies.
New MCMC method learns sparse preconditioner for high-dimensional problems.
New method for estimating financial covariance matrices efficiently.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
This paper improves linear system solving by optimizing matrix diagonal scaling.
On the level of Lie algebras, the contraction procedure is a method to create a new Lie algebra from a given Lie algebra by rescaling generators and letting the scaling parameter tend to zero. One of the most well-known examples is the contraction from su(2) to e(2), the Lie algebra of upper-triangular matrices with ze…
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating why this algorithm lacks in robustness then reveals two interesting insights. The first insight suggests a new way to scale the stepsizes, …