This paper solves matrix blind joint block diagonalization with noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…
Develops a novel stochastic algorithm for diagonal estimation of large matrices.
This paper optimizes diagonal preconditioning to improve matrix condition numbers.
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
Localized sketching improves matrix multiplication and ridge regression complexity.
New insights into Hessian structure of neural networks reveal two forces.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
In (exploratory) factor analysis, the loading matrix is identified only up to orthogonal rotation. For identifiability, one thus often takes the loading matrix to be lower triangular with positive diagonal entries. In Bayesian inference, a standard practice is then to specify a prior under which the loadings are indepe…
Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.
Two Fisher information matrix estimators are analyzed for neural networks, focusing on their variances and trade-offs.
We provide an online convex optimization algorithm with regret that interpolates between the regret of an algorithm using an optimal preconditioning matrix and one using a diagonal preconditioning matrix. Our regret bound is never worse than that obtained by diagonal preconditioning, and in certain setting even surpass…
Adaptive stochastic gradient methods such as AdaGrad have gained popularity in particular for training deep neural networks. The most commonly used and studied variant maintains a diagonal matrix approximation to second order information by accumulating past gradients which are used to tune the step size adaptively. In…
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
The adaptive gradient online learning method known as AdaGrad has seen widespread use in the machine learning community in stochastic and adversarial online learning problems and more recently in deep learning methods. The method's full-matrix incarnation offers much better theoretical guarantees and potentially better…
New model reduces matrix factorization bias, yielding truly low-rank solutions.
Gaussian graphical models are widely utilized to infer and visualize networks of dependencies between continuous variables. However, inferring the graph is difficult when the sample size is small compared to the number of variables. To reduce the number of parameters to estimate in the model, we propose a non-asymptoti…
A Semi-Hidden Markov Model (SHMM) for bursty error channels is defined by a state transition probability matrix , a prior probability vector , and the state dependent output symbol error probability matrix . Several processes are utilized for estimating , and from a given empirically obtained or sim…
A new method solves diagonally constrained SDPs quickly and accurately.
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
Homogeneous links were introduced by Peter Cromwell, who proved that the projection surface of these links, that given by the Seifert algorithm, has minimal genus. Here we provide a different proof, with a geometric rather than combinatorial flavor. To do this, we first show a direct relation between the Seifert matrix…
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
A new metric learning framework for signed graphs using Gershgorin disc alignment.
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank represent…
Efficiently approximates Sparse PCA with significant speedups and minor error.
Variational Bayesian neural networks combine the flexibility of deep learning with Bayesian uncertainty estimation. However, inference procedures for flexible variational posteriors are computationally expensive. A recently proposed method, noisy natural gradient, is a surprisingly simple method to fit expressive poste…
Novel risk matrix for optimal portfolio choice with tail risk considerations.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
Method estimates M-matrices in graphical models with improved accuracy.
The paper tackles sparse graph learning under Laplacian-related constraints, improving upon existing methods.
Randomized block-diagonal preconditioning improves parallel learning convergence.
This paper improves linear system solving by optimizing matrix diagonal scaling.
New method estimates sparse covariance matrices in logit mixtures.
Shrunk sample covariance matrix is a factor model of a special form combining some (typically, style) risk factor(s) and principal components with a (block-)diagonal factor covariance matrix. As such, shrinkage, which essentially inherits out-of-sample instabilities of the sample covariance matrix, is not an alternativ…
New MCMC method learns sparse preconditioner for high-dimensional problems.
Recurrent neural networks (RNNs) have been successfully used on a wide range of sequential data problems. A well known difficulty in using RNNs is the \textit{vanishing or exploding gradient} problem. Recently, there have been several different RNN architectures that try to mitigate this issue by maintaining an orthogo…
The crossing matrix of a braid on strands is the integer matrix with zero diagonal whose entry is the algebraic number (positive minus negative) of crossings by strand over strand . When restricted to the subgroup of pure braids, this defines a homomorphism onto the additive subgroup of $N…
A new model captures multifractal volatility in stock returns.
Second-order methods for neural network optimization have several advantages over methods based on first-order gradient descent, including better scaling to large mini-batch sizes and fewer updates needed for convergence. But they are rarely applied to deep learning in practice because of high computational cost and th…
A new model captures multifractal volatility in stock returns.
The paper identifies redundant columns in matrices for feature selection and clustering.
We discuss a clustering method for Gaussian mixture model based on the sparse principal component analysis (SPCA) method and compare it with the IF-PCA method. We also discuss the dependent case where the covariance matrix is not necessarily diagonal.
A novel tracking algorithm models dynamic objects as ellipsoids with time-varying orientation.
New methods improve solving linear systems and preconditioning with reduced complexity.
Study on nilpotent Lie algebras with specific metrics.
An algorithm for computing positive semidefinite factorizations of matrices.
New algorithms estimate matrix norms without matrix multiplication.