This paper solves matrix blind joint block diagonalization with noise.
problem Identifying the diagonalizer and block diagonal structure of matrices under noise.
method Bi-block diagonalization method.
result The method can identify the exact solution under certain conditions.
New adaptive methods improve deep learning performance.
problem Training deep networks efficiently and effectively.
method Block-diagonal matrix adaptation for gradient updates.
result Block-diagonal methods outperform adaptive diagonal methods and vanilla SGD.
Develops a novel stochastic algorithm for diagonal estimation of large matrices.
problem Efficient diagonal estimation for large or implicit matrices.
method Adaptive parameter selection in a stochastic algorithm.
result Lower bound on random query vectors needed for estimation.
This paper optimizes diagonal preconditioning to improve matrix condition numbers.
problem Optimizing diagonal preconditioning to reduce matrix condition numbers.
method Reformulated as a quasi-convex problem, solved with bisection and Newton updates.
result Optimal diagonal preconditioners can significantly improve iterative methods.
Constructs coordinates to diagonalize Toda flow on matrices with simple spectrum.
problem Diagonalizing the Toda flow on matrices with simple spectrum.
method Lie theoretic methods applied to complex semisimple Lie algebras and their real forms.
result Decouples the Toda vector field into simpler components.
Localized sketching improves matrix multiplication and ridge regression complexity.
problem Efficiently approximate matrix multiplication and ridge regression with limited data availability.
method Localized sketching matrices for block diagonal structure, reducing sample complexity.
result Localized sketching achieves sample complexity matching global sketching methods.
New insights into Hessian structure of neural networks reveal two forces.
problem Understanding the Hessian structure of neural networks.
method Analyzing the static and dynamic forces, comparing limit distributions using random matrix theory.
result The Hessian structure arises from a combination of static and dynamic forces, with C being a primary driver. In (exploratory) factor analysis, the loading matrix is identified only up to orthogonal rotation. For identifiability, one thus often takes the loading matrix to be lower triangular with positive diagonal entries. In Bayesian inference, a standard practice is then to specify a prior under which the loadings are indepe…
The paper simplifies the Fisher information matrix for random deep networks, speeding up learning.
problem Learning deep neural networks efficiently with large parameter spaces.
method Statistical neurodynamical method to reveal Fisher information properties, proving unit-wise block diagonal structure and explicit inverse.
result Explicit natural gradient formula without matrix inversion, speeding up learning.
Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.
problem Subspace clustering with block diagonal structure for noisy data.
method ABDR explicitly pursues block diagonality without sacrificing convexity, using a specially designed convex regularizer.
result Experimental results show ABDR outperforms state-of-the-arts.
Modular method simplifies curvature computation in neural nets.
problem Efficient computation of curvature matrices for training neural nets.
method Modular backpropagation for block-diagonal approximations.
result Compact notation and easy integration into machine learning libraries.
Two Fisher information matrix estimators are analyzed for neural networks, focusing on their variances and trade-offs.
problem Estimating the Fisher information matrix in neural networks due to its high computational cost.
method Examined two popular diagonal Fisher information matrix estimators and their variances in neural networks for regression and classification.
result The variances of the estimators depend on the non-linearity with respect to different parameter groups and should not be neglected.
A method to simplify Gaussian graphical models for high-dimensional data.
problem Difficulty in inferring networks of dependencies between variables when sample size is small.
method Approximate covariance matrix as block-diagonal, select threshold using slope heuristic, infer network in blocks.
result The method reduces the number of parameters to estimate and improves network inference.
Improved online learning algorithm with better regret bounds.
problem Online convex optimization with improved regret bounds.
method Matrix-free preconditioning approach.
result Our algorithm outperforms diagonal preconditioning in certain settings.
In this paper we show that the matrix of chromatic joins and the Gram matrix of the Temperley-Lieb algebra are similar (after rescaling), with the change of basis given by diagonal matrices.
Efficient approximations for AdaGrad reduce computation while maintaining performance.
problem Training deep neural networks efficiently in high dimensions.
method Ada-LR and RadaGrad use random projections to approximate full-matrix AdaGrad.
result Regret of Ada-LR is close to full-matrix AdaGrad, achieving similar performance with less computation.
New model reduces matrix factorization bias, yielding truly low-rank solutions.
problem Gradient descent's implicit bias in matrix factorization.
method Introducing a new factorization model with constrained factors and diagonal components.
result The new model consistently exhibits a strong implicit bias, yielding truly low-rank solutions.
Improved Hessian-free method for neural networks reduces computational cost.
problem High computational cost and model-dependent algorithmic variations in second-order methods.
method Block-diagonal approximation of the generalized Gauss-Newton matrix, conjugate gradient updates for each block.
result Better convergence and generalization compared to original Hessian-free and Adam methods.
A new method solves diagonally constrained SDPs quickly and accurately.
problem Solving large-scale diagonally constrained SDPs efficiently.
method Combines momentum from convex optimization with coordinate descent and matrix factorization.
result Local linear convergence and first-order critical point convergence proved.
CompAdaGrad improves AdaGrad's performance without its computational cost.
problem Improving AdaGrad's performance without its high computational cost.
method CompAdaGrad combines full-matrix and diagonal regularization in a low-dimensional subspace.
result CompAdaGrad achieves better results than diagonal AdaGrad with linear computational complexity.
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
problem Scaling Gaussian processes for large datasets.
method Structured diagonal scaling matrix and Power-EP framework.
result Structured approximations improve performance without increasing computational cost.
Generates correlation matrices with specific graph structures using convex optimization.
problem Creating theoretical correlation matrices with prescribed graph structures.
method Convex optimization framework projecting an initial matrix onto an elliptope with positive semidefiniteness constraint.
result The approach offers greater flexibility in generating correlation matrices with controlled mean of off-diagonal entries.
Homogeneous links were introduced by Peter Cromwell, who proved that the projection surface of these links, that given by the Seifert algorithm, has minimal genus. Here we provide a different proof, with a geometric rather than combinatorial flavor. To do this, we first show a direct relation between the Seifert matrix…
Sparse PCA method for clustering Gaussian mixtures.
problem Clustering Gaussian mixture models.
method Sparse Principal Component Analysis (SPCA) for clustering.
result Comparison with IF-PCA method and discussion of non-diagonal covariance matrices.
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
problem Minimizing sharpness in diagonal linear networks.
method Stochastic sharpness-aware minimization (SAM) with isotropic noise.
result Noise forces shrinkage-thresholding of true parameters.
Paper proposes a new Markov model for efficient PLC system design.
problem Efficient estimation of Markov model parameters for bursty error channels.
method Introduced a Block Diagonal Markov model and a modified Baum-Welch algorithm.
result Efficient estimation of state transition matrix Λ for PLC system design. Paper proposes an effective mean-field inference method for NNBMs.
problem Inference in NNBMs is challenging due to their complex structure.
method Uses mean-field method and diagonal consistency method.
result Effective inference method for NNBMs is proposed.
A new metric learning framework for signed graphs using Gershgorin disc alignment.
problem Learning Mahalanobis metrics from signed graphs efficiently.
method Proposes a fast metric learning framework using Gershgorin disc perfect alignment (GDPA) to circumvent full eigen-decomposition.
result Proves that Gershgorin disc left-ends of similarity transform are perfectly aligned at the smallest eigenvalue, enabling efficient optimization.
Efficiently approximates Sparse PCA with significant speedups and minor error.
problem Sparse Principal Component Analysis (Sparse PCA) is NP-hard and computationally expensive.
method Approximates the covariance matrix with block-diagonal form, solves sub-problems in each block, and reconstructs the solution.
result Significant computational speedups with minor additive error.
Improved neural network inference with eigenvalue correction.
problem Inference of flexible variational posteriors is computationally expensive.
method Eigenvalue correction to matrix-variate Gaussian posterior.
result Empirically, the method outperforms existing algorithms.
We propose an efficient method for approximating natural gradient descent in neural networks which we call Kronecker-Factored Approximate Curvature (K-FAC). K-FAC is based on an efficiently invertible approximation of a neural network's Fisher information matrix which is neither diagonal nor low-rank, and in some cases…
Novel risk matrix for optimal portfolio choice with tail risk considerations.
problem Optimal portfolio choice with tail risk events.
method Risk matrix with Value-at-Risk and Delta-CoVaR measures, derived conditions for closed-form solution, examination of portfolio risk and centrality, demonstration of asset centrality's impact on optimal weight allocation.
result Portfolio risk is not necessarily increasing with stock centrality and can be improved by high connectivity.
Efficient subspace clustering using Kronecker product reduces computational complexity.
problem Efficiency and scalability issues in traditional subspace clustering methods for large datasets.
method Proposes a subspace clustering model based on the Kronecker product to reduce computational complexity.
result Significantly improved efficiency compared to state-of-the-art methods on public datasets.
Method estimates M-matrices in graphical models with improved accuracy.
problem Estimating M-matrices as precision matrices in Gaussian graphical models.
method Adaptive multiple-stage estimation method solving weighted ℓ1-regularized problems.
result Method outperforms state-of-the-art methods in precision matrix estimation and graph edge identification.
This paper improves linear system solving by optimizing matrix diagonal scaling.
problem Improving the condition number of a matrix for faster iterative methods.
method Left or right diagonal rescaling of the matrix A, with new bounds and algorithms.
result Jacobi preconditioning reduces A's condition number to within a quadratic factor of the best possible scaling.
The paper tackles sparse graph learning under Laplacian-related constraints, improving upon existing methods.
problem Learning a sparse undirected graph from multivariate data under Laplacian-related constraints.
method Modifications to penalized log-likelihood approaches to enforce total positivity and lasso/adaptive lasso penalties using ADMM.
result The proposed constrained adaptive lasso approach significantly outperforms existing Laplacian-based approaches.
Randomized block-diagonal preconditioning improves parallel learning convergence.
problem Improving convergence of gradient-based optimization methods in parallel settings.
method Randomization of coordinates during optimization to repartition tasks.
result Randomization significantly improves convergence of block-diagonal preconditioned methods.
New method estimates sparse covariance matrices in logit mixtures.
problem Estimating correlations among random coefficients in logit models.
method Mixed-integer optimization (MIO) with Markov Chain Monte Carlo (MCMC) for posterior draws.
result Correctly recovers true covariance structure from synthetic data.
New MCMC method learns sparse preconditioner for high-dimensional problems.
problem High-dimensional sampling with complex correlation structures.
method Adaptive MCMC with sparse preconditioner using online PCA.
result Significant reduction in computational complexity and improved performance.
A new model captures multifractal volatility in stock returns.
problem Capturing multifractal volatility in stock returns.
method Introduced mLog S-fBM model, defined mS-fBM, and developed calibration procedure.
result Model captures multifractal behavior in stock returns, validating on real data.
Paper proposes a new method for brain disease classification using connectome data.
problem Challenges in classifying brain diseases due to small sample size and high dimensionality.
method Simultaneous approximate diagonalization of adjacency matrices to compute stable eigenstructures.
result The method outperforms simple baselines and state-of-the-art approaches for Alzheimer's disease detection.
A new model captures multifractal volatility in stock returns.
problem Capturing multifractal volatility in stock returns.
method Introduced mLog S-fBM model, defined mS-fBM, and developed calibration procedure.
result Validated model on synthetic and real data, showing multifractal behavior.
Efficient neural networks compute various differential operators cheaply.
problem Efficient computation of higher time complexity differential operators.
method Restricted neural network architectures with diagonal and hollow Jacobian matrices, allowing efficient extraction of dimension-wise derivatives.
result Demonstrated efficient computation of differential operators for various applications.
New model complexes help solve homology problems in 4D space.
problem Computing homology of 2-complexes in 4D space.
method Use model 2-complexes built from group presentations, showing linear embeddability.
result Homology computation is equivalent to matrix diagonalization in 4D space.
A novel tracking algorithm models dynamic objects as ellipsoids with time-varying orientation.
problem Tracking dynamic objects with time-varying orientation.
method Random matrix framework with variational Bayes for non-linear inference.
result The method outperforms state-of-the-art methods in accuracy and robustness.
The paper identifies redundant columns in matrices for feature selection and clustering.
problem Identifying redundant columns in matrices for feature selection and clustering.
method Proves that after re-ordering columns, a matrix can be block-diagonalized revealing linearly dependent columns.
result Identifies redundant columns in matrices, aiding in feature selection and clustering.
New methods improve solving linear systems and preconditioning with reduced complexity.
problem Efficiently solving linear systems and preconditioning matrices.
method Developed structured semidefinite programming algorithms.
result Improved runtimes for preconditioning and solving linear systems.
An algorithm for computing positive semidefinite factorizations of matrices.
problem Computing positive semidefinite factorizations of matrices.
method Non-commutative extension of Lee-Seung's algorithm (Matrix Multiplicative Update, MMU).
result The MMU algorithm ensures PSD updates and achieves critical points.