New method improves Pham's algorithm for joint diagonalization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved spectral methods of moments for robust latent variable model learning.
This paper solves matrix blind joint block diagonalization with noise.
Framework for incomplete multi-view learning improves efficiency and clustering accuracy.
We consider moment matching techniques for estimation in Latent Dirichlet Allocation (LDA). By drawing explicit links between LDA and discrete versions of independent component analysis (ICA), we first derive a new set of cumulant-based tensors, with an improved sample complexity. Moreover, we reuse standard ICA techni…
We explore the connection between two problems that have arisen independently in the signal processing and related fields: the estimation of the geometric mean of a set of symmetric positive definite (SPD) matrices and their approximate joint diagonalization (AJD). Today there is a considerable interest in estimating t…
Non-orthogonal joint diagonalization (NJD) free of prewhitening has been widely studied in the context of blind source separation (BSS) and array signal processing, etc. However, NJD is used to retrieve the jointly diagonalizable structure for a single set of target matrices which are mostly formulized with a single da…
The paper introduces a new Wasserstein distance for approximating posteriors in inverse problems.
Recently, there has been a trend to combine independent component analysis and canonical polyadic decomposition (ICA-CPD) for an enhanced robustness for the computation of CPD, and ICA-CPD could be further converted into CPD of a 5th-order partially symmetric tensor, by calculating the eigenmatrices of the 4th-order cu…
Modular method simplifies curvature computation in neural nets.
We introduce three novel semi-parametric extensions of probabilistic canonical correlation analysis with identifiability guarantees. We consider moment matching techniques for estimation in these models. For that, by drawing explicit links between the new models and a discrete version of independent component analysis …
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
Optimization algorithms that leverage gradient covariance information, such as variants of natural gradient descent (Amari, 1998), offer the prospect of yielding more effective descent directions. For models with many parameters, the covariance matrix they are based on becomes gigantic, making them inapplicable in thei…
Diagonal linear networks converge to lasso regularization path during training.
Efficiently approximates Sparse PCA with significant speedups and minor error.
Proposes a method to predict responses from covariates over time.
Develops large-sample theory for non-stationary source separation.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
SGD on diagonal linear networks approximates to SDE in high dimensions.
Apollo improves nonconvex stochastic optimization efficiency.
New method improves deep learning model robustness and accuracy for long sequences.
Joint learning framework for clustering and graph construction.
We consider the problem of approximate joint triangularization of a set of noisy jointly diagonalizable real matrices. Approximate joint triangularizers are commonly used in the estimation of the joint eigenstructure of a set of matrices, with applications in signal processing, linear algebra, and tensor decomposition.…
Diagonal transformations preserve independence structures in non-Gaussian distributions.
This paper considers the problem of brain disease classification based on connectome data. A connectome is a network representation of a human brain. The typical connectome classification problem is very challenging because of the small sample size and high dimensionality of the data. We propose to use simultaneous app…
Unified theorem for deep and shallow joint-equivariant machines.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
We study the perturbations of two classes of static black ellipsoid solutions of four dimensional vacuum Einstein equations. Such solutions are described by generic off--diagonal metrics which are generated by anholonomic transforms of diagonal metrics. The analysis is performed in the approximation of small eccentrici…
SLANG improves uncertainty estimation in deep learning models.
New model reduces matrix factorization bias, yielding truly low-rank solutions.
A susceptibility propagation that is constructed by combining a belief propagation and a linear response method is used for approximate computation for Markov random fields. Herein, we formulate a new, improved susceptibility propagation by using the concept of a diagonal matching method that is based on mean-field app…
Adaptive stochastic gradient methods such as AdaGrad have gained popularity in particular for training deep neural networks. The most commonly used and studied variant maintains a diagonal matrix approximation to second order information by accumulating past gradients which are used to tune the step size adaptively. In…
Sharp results link DLN gradient flow to basis pursuit optimization and GHA phase transitions.
Develops efficient quasi-Newton methods for training deep neural networks.
Adler had shown in 1979 that the Toda system can be given a coad- joint orbit description. We quantize the Toda system by viewing it as a single orbit of a multiplicative group of lower triangular matrices of determinant one with pos- itive diagonal entries. We get a unitary representation of the group with square inte…
A fair PCA method using JEVD ensures balanced data representation.
Second-order methods for neural network optimization have several advantages over methods based on first-order gradient descent, including better scaling to large mini-batch sizes and fewer updates needed for convergence. But they are rarely applied to deep learning in practice because of high computational cost and th…
Localized sketching improves matrix multiplication and ridge regression complexity.
Improved multimodal variational models capture more complex joint distributions.
We collect well known and less known facts about the bivariate normal distribution and translate them into copula language. In addition, we prove a very general formula for the bivariate normal copula, we compute Gini's gamma, and we provide improved bounds and approximations on the diagonal.
We consider Bayesian inference problems with computationally intensive likelihood functions. We propose a Gaussian process (GP) based method to approximate the joint distribution of the unknown parameters and the data. In particular, we write the joint density approximately as a product of an approximate posterior dens…
SEM-DNN learns reciprocal interactions from observational data without external instruments.
Study approximates top Lyapunov exponents for surface mapping classes.
This study explains why approximate NGD works well in wide neural networks.
This work improves OOD detection using deep generative models by approximating Fisher information metrics.
In the first quarter of 2006 Chicago Board Options Exchange (CBOE) introduced, as one of the listed products, options on its implied volatility index (VIX). This created the challenge of developing a pricing framework that can simultaneously handle European options, forward-starts, options on the realized variance and …
The Bethe free energy approximation is reliable when convex on a submanifold, the 'Bethe box'.
We introduce a general framework for estimation of inverse covariance, or precision, matrices from heterogeneous populations. The proposed framework uses a Laplacian shrinkage penalty to encourage similarity among estimates from disparate, but related, subpopulations, while allowing for differences among matrices. We p…