In this paper, we study deep diagonal circulant neural networks, that is deep neural networks in which weight matrices are the product of diagonal and circulant ones. Besides making a theoretical analysis of their expressivity, we introduced principled techniques for training these models: we devise an initialization s…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper proves a distribution claim for neural network Jacobians.
Method estimates M-matrices in graphical models with improved accuracy.
A fast metric learning framework using Gershgorin disc alignment.
Gradient flow on softmax attention minimizes nuclear norm of weight matrices.
Diagonal linear networks converge to lasso regularization path during training.
Diagonal transformations preserve independence structures in non-Gaussian distributions.
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
Independent Component Analysis (ICA) - one of the basic tools in data analysis - aims to find a coordinate system in which the components of the data are independent. In this paper we present Multiple-weighted Independent Component Analysis (MWeICA) algorithm, a new ICA method which is based on approximate diagonalizat…
The approximate joint diagonalization of a set of matrices consists in finding a basis in which these matrices are as diagonal as possible. This problem naturally appears in several statistical learning tasks such as blind signal separation. We consider the diagonalization criterion studied in a seminal paper by Pham (…
Develops a novel stochastic algorithm for diagonal estimation of large matrices.
Three methods for tuning HMC diagonal scale matrices compared.
In this paper, we propose a new Recurrent Neural Network (RNN) architecture. The novelty is simple: We use diagonal recurrent matrices instead of full. This results in better test likelihood and faster convergence compared to regular full RNNs in most of our experiments. We show the benefits of using diagonal recurrent…
Recurrent Neural Networks (RNNs) are designed to handle sequential data but suffer from vanishing or exploding gradients. Recent work on Unitary Recurrent Neural Networks (uRNNs) have been used to address this issue and in some cases, exceed the capabilities of Long Short-Term Memory networks (LSTMs). We propose a simp…
This paper solves matrix blind joint block diagonalization with noise.
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
Paper estimates GMMs with unknown covariances using sparse regularization.
Gradient descent optimally trains RNNs without overparameterization.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
Algorithm solves robust linear regression with block Lewis weights.
We propose a novel approach to addressing the vanishing (or exploding) gradient problem in deep neural networks. We construct a new architecture for deep neural networks where all layers (except the output layer) of the network are a combination of rotation, permutation, diagonal, and activation sublayers which are all…
This paper optimizes diagonal preconditioning to improve matrix condition numbers.
The crossing matrix of a braid on strands is the integer matrix with zero diagonal whose entry is the algebraic number (positive minus negative) of crossings by strand over strand . When restricted to the subgroup of pure braids, this defines a homomorphism onto the additive subgroup of $N…
Localized sketching improves matrix multiplication and ridge regression complexity.
Non-orthogonal joint diagonalization (NJD) free of prewhitening has been widely studied in the context of blind source separation (BSS) and array signal processing, etc. However, NJD is used to retrieve the jointly diagonalizable structure for a single set of target matrices which are mostly formulized with a single da…
The paper explores spinors and polyforms using quaternions and octonions.
We study algebraic properties of matrices whose rows are mutual neighbours, and are also neigbours of 0 ("neighbour" in the sense of a certain nilpotency condition). The intended application is in synthetic differential geometry. For a square matrix of this kind, the product of the diagonal entries equals the determina…
This work is motivated by numerical solutions to Hamilton-Jacobi-Bellman quasi-variational inequalities (HJBQVIs) associated with combined stochastic and impulse control problems. In particular, we consider (i) direct control, (ii) penalized, and (iii) semi-Lagrangian discretization schemes applied to the HJBQVI proble…
Study on likelihood functions, associative equations, and Frobenius manifolds.
The high-order relations between the content in social media sharing platforms are frequently modeled by a hypergraph. Either hypergraph Laplacian matrix or the adjacency matrix is a big matrix. Randomized algorithms are used for low-rank factorizations in order to approximately decompose and eventually invert such big…
Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…
We prove ultradifferentiable Chevelley restriction theorems for a wide range of ultradifferentiable classes. As a special case we find that isotropic functions, i.e., functions defined on the vector space of real symmetric matrices invariant under the action of the special orthogonal group by conjugation, possess some …
Researchers develop geodesics for a new metric on correlation matrices.
This note classifies splittable lattices in a specific Lie group.
We present explicit formulas for the coordinates in which the Hamiltonians of the Benenti systems with flat metrics take natural form and the metrics in question are represented by constant diagonal matrices.
A new metric learning framework for signed graphs using Gershgorin disc alignment.
In this paper we consider the use of the space vs. time Kronecker product decomposition in the estimation of covariance matrices for spatio-temporal data. This decomposition imposes lower dimensional structure on the estimated covariance matrix, thus reducing the number of samples required for estimation. To allow a sm…
We give an overview of the generalized Calderón-Zygmund theory for "non-integral" singular operators, that is, operators without kernels bounds but appropriate off-diagonal estimates. This theory is powerful enough to obtain weighted estimates for such operators and their commutators with $\BMO$ functions. of…
Despite their successes, what makes kernel methods difficult to use in many large scale problems is the fact that storing and computing the decision function is typically expensive, especially at prediction time. In this paper, we overcome this difficulty by proposing Fastfood, an approximation that accelerates such co…
The paper develops efficient algorithms for variational inference with mixtures of isotropic Gaussians.
Embedding complex objects as vectors in low dimensional spaces is a longstanding problem in machine learning. We propose in this work an extension of that approach, which consists in embedding objects as elliptical probability distributions, namely distributions whose densities have elliptical level sets. We endow thes…
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
The paper models financial correlation matrices using permutation invariant Gaussian models and predicts market anomalies.
New MCMC method learns sparse preconditioner for high-dimensional problems.
An algorithm for computing positive semidefinite factorizations of matrices.
New method for estimating financial covariance matrices efficiently.