Diagonal transformations preserve independence structures in non-Gaussian distributions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
The approximate joint diagonalization of a set of matrices consists in finding a basis in which these matrices are as diagonal as possible. This problem naturally appears in several statistical learning tasks such as blind signal separation. We consider the diagonalization criterion studied in a seminal paper by Pham (…
Develops a novel stochastic algorithm for diagonal estimation of large matrices.
Three methods for tuning HMC diagonal scale matrices compared.
Method estimates M-matrices in graphical models with improved accuracy.
In this paper, we propose a new Recurrent Neural Network (RNN) architecture. The novelty is simple: We use diagonal recurrent matrices instead of full. This results in better test likelihood and faster convergence compared to regular full RNNs in most of our experiments. We show the benefits of using diagonal recurrent…
This paper solves matrix blind joint block diagonalization with noise.
A fast metric learning framework using Gershgorin disc alignment.
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
In this paper, we study deep diagonal circulant neural networks, that is deep neural networks in which weight matrices are the product of diagonal and circulant ones. Besides making a theoretical analysis of their expressivity, we introduced principled techniques for training these models: we devise an initialization s…
This paper optimizes diagonal preconditioning to improve matrix condition numbers.
The crossing matrix of a braid on strands is the integer matrix with zero diagonal whose entry is the algebraic number (positive minus negative) of crossings by strand over strand . When restricted to the subgroup of pure braids, this defines a homomorphism onto the additive subgroup of $N…
Diagonal linear networks converge to lasso regularization path during training.
Localized sketching improves matrix multiplication and ridge regression complexity.
Non-orthogonal joint diagonalization (NJD) free of prewhitening has been widely studied in the context of blind source separation (BSS) and array signal processing, etc. However, NJD is used to retrieve the jointly diagonalizable structure for a single set of target matrices which are mostly formulized with a single da…
The paper explores spinors and polyforms using quaternions and octonions.
The paper proves a distribution claim for neural network Jacobians.
We study algebraic properties of matrices whose rows are mutual neighbours, and are also neigbours of 0 ("neighbour" in the sense of a certain nilpotency condition). The intended application is in synthetic differential geometry. For a square matrix of this kind, the product of the diagonal entries equals the determina…
This work is motivated by numerical solutions to Hamilton-Jacobi-Bellman quasi-variational inequalities (HJBQVIs) associated with combined stochastic and impulse control problems. In particular, we consider (i) direct control, (ii) penalized, and (iii) semi-Lagrangian discretization schemes applied to the HJBQVI proble…
Study on likelihood functions, associative equations, and Frobenius manifolds.
Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…
Researchers develop geodesics for a new metric on correlation matrices.
This note classifies splittable lattices in a specific Lie group.
We present explicit formulas for the coordinates in which the Hamiltonians of the Benenti systems with flat metrics take natural form and the metrics in question are represented by constant diagonal matrices.
A new metric learning framework for signed graphs using Gershgorin disc alignment.
Chevalley theorems extended to isotropic functions on matrix spaces.
In this paper we consider the use of the space vs. time Kronecker product decomposition in the estimation of covariance matrices for spatio-temporal data. This decomposition imposes lower dimensional structure on the estimated covariance matrix, thus reducing the number of samples required for estimation. To allow a sm…
Despite their successes, what makes kernel methods difficult to use in many large scale problems is the fact that storing and computing the decision function is typically expensive, especially at prediction time. In this paper, we overcome this difficulty by proposing Fastfood, an approximation that accelerates such co…
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
New method estimates sparse covariance matrices in logit mixtures.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
The paper models financial correlation matrices using permutation invariant Gaussian models and predicts market anomalies.
New MCMC method learns sparse preconditioner for high-dimensional problems.
An algorithm for computing positive semidefinite factorizations of matrices.
New method for estimating financial covariance matrices efficiently.
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
New nodal domain theorems for symmetric matrices via signed graphs.
This paper considers the problem of brain disease classification based on connectome data. A connectome is a network representation of a human brain. The typical connectome classification problem is very challenging because of the small sample size and high dimensionality of the data. We propose to use simultaneous app…
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
Paper estimates GMMs with unknown covariances using sparse regularization.
New method improves deep learning model robustness and accuracy for long sequences.
We explore the connection between two problems that have arisen independently in the signal processing and related fields: the estimation of the geometric mean of a set of symmetric positive definite (SPD) matrices and their approximate joint diagonalization (AJD). Today there is a considerable interest in estimating t…
New model handles complex non-linear relationships with hidden graph structures.
Study of correlated Wigner matrices with BBP transitions.
The inverse covariance matrix provides considerable insight for understanding statistical models in the multivariate setting. In particular, when the distribution over variables is assumed to be multivariate normal, the sparsity pattern in the inverse covariance matrix, commonly referred to as the precision matrix, cor…
New estimators reduce computation for Kendall's tau and conditional Kendall's tau matrices under structural assumptions.