New bound for neural networks with full-rank weights, independent of network width.
problem Understanding generalization of neural networks with full-rank weight matrices.
method Using Koopman operators to derive a tighter generalization bound for full-rank weight matrices.
result The bound is tighter than existing norm-based bounds when condition numbers are small.
New metrics defined for full-rank correlation matrices, ensuring unique operations.
problem No suitable problem statement as the abstract does not describe a problem to be solved.
method New Riemannian metrics defined on full-rank correlation matrices, providing unique operations.
result Unique Riemannian logarithm and Fréchet mean defined for full-rank correlation matrices.
A fast method for multichannel source separation using jointly diagonalizable SCMs.
problem Computational inefficiency and poor performance in multichannel source separation.
method Restricts SCMs to jointly-diagonalizable but full-rank matrices, proposing efficient algorithms.
result Significant speedup and improved performance compared to original methods.
New method improves matrix factorization accuracy and speed.
problem Improving matrix factorization for large, noisy data.
method Introducing generalized round-rank (GRR) for ordinal-valued matrices.
result GRR-based matrices cannot be well approximated by low-rank linear factorization.
Low-rank matrices explain data science patterns.
problem Why do data matrices often have low rank?
method A generative model with latent variables and piecewise functions.
result Approximating large matrices with low rank is feasible.
A new method for efficiently updating large-scale matrices in real-time.
problem Updating large-scale matrices with evolving data in real-time.
method Incremental SVD approach that handles row/column appends, rank-1 updates, and refresh strategies.
result Incremental SVD achieves accuracy close to full SVD with a fraction of the computational cost.
Improves BBVI for high-dimensional Gaussian approximations by using low-rank approximations.
problem Scalability issues with BBVI for high-dimensional multivariate Gaussian approximations.
method Extends BaM framework to handle full covariance matrices by integrating patch step for low-rank parameterization.
result Shows improved efficiency and scalability on synthetic and real-world high-dimensional inference problems.
New framework estimates eigenvalues of kernel matrices without full matrix construction.
problem Estimating eigenvalues of large kernel matrices efficiently.
method Eigenvalue quantile estimation framework for kernel matrices with quick decay.
result Validates framework with empirical evidence and proves interlacing theorem.
Mirror descent algorithm recovers low-rank matrices in matrix sensing.
problem Matrix sensing with low-rank matrices under certain conditions.
method Discrete-time mirror descent applied to empirical risk with Bregman divergence analysis.
result Mirror descent converges to a matrix minimizing a specific nuclear norm-related quantity.
Efficiently implements MEG for low-rank matrix optimization problems.
problem Optimization over spectrahedron with low-rank matrices.
method Matrix Exponentiated Gradient (MEG) method with efficient implementations.
result Methods converge from a warm-start initialization with similar rates to full-SVD-based counterparts.
Efficient and accurate low-rank approximations of multiple data sources are essential in the era of big data. The scaling of kernel-based learning algorithms to large datasets is limited by the O(n^2) computation and storage complexity of the full kernel matrix, which is required by most of the recent kernel learning a…
New method corrects quantization errors in LLMs using low-rank matrices.
problem Correcting quantization errors in large language models.
method Introducing low-rank weight matrices to correct quantized activations in LLMs.
result Reduces accuracy gap with original model by more than 50% using low-rank matrices.
New method for initializing low-rank neural networks improves performance.
problem Training low-rank neural networks efficiently and accurately.
method Inspired by function approximation, proposes a novel low-rank initialization framework.
result Demonstrates significant gap between spectral and low-rank initialization approaches.
Researchers develop geodesics for a new metric on correlation matrices.
problem Lack of intrinsic tools for statistical analyses of correlation matrices.
method Developed geodesics for the quotient-affine metric on full-rank correlation matrices.
result Provided fundamental Riemannian operations for the quotient-affine metric.
RKPCA improves robustness of PCA for high-rank matrices.
problem Robust recovery of high-rank matrices corrupted by sparse noises.
method RKPCA decomposes matrices into sparse and low-rank components.
result RKPCA provides high recovery accuracy with theoretical guarantees.
In this paper, we develop a relative error bound for nuclear norm regularized matrix completion, with the focus on the completion of full-rank matrices. Under the assumption that the top eigenspaces of the target matrix are incoherent, we derive a relative upper bound for recovering the best low-rank approximation of t…
New insights into attention mechanisms reveal dramatic trade-offs between rank and heads.
problem Dramatic trade-offs between rank and number of heads in attention mechanisms.
method Presented a simple target function and proved theoretical limits.
result Full-rank attention is necessary for long contexts, while low-rank is sufficient for short ones.
The paper explains geometrically why certain mappings have singular points.
problem Understanding singular points in mappings from R^2 to R^3 and higher.
method Analyzing full rank matrices constructed from coefficients of mappings.
result Mappings have only one singular point when ℓ=3 and no singular points when ℓ>3.
Article presents QR and LQ decomposition algorithms for various matrix sizes and ranks.
problem Solving least squares problems in machine learning and computer vision.
method Developed novel matrix backpropagation algorithms for QR and LQ decompositions of different matrix sizes and ranks.
result Numerical stability and computational efficiency of the proposed methods.
This paper uses rank correlation methods to construct MSTs from financial returns, finding them more stable and robust.
problem Stability and robustness of MSTs constructed from financial correlation matrices.
method Pearson, Spearman, and Kendall's τ rank correlation methods applied to daily financial returns. result Rank MSTs are more stable and robust than MSTs constructed using Pearson correlation.
New framework uses symmetry-based matrices for efficient, flexible NNs.
problem Designing neural networks with relaxed equivariance.
method Symmetry-based structured matrices, Group Matrices (GMs).
result GMs enable competitive performance with fewer parameters.
New methods extend kernel estimators for partial rankings, improving performance in machine learning tasks.
problem Incomplete rankings data in real-world applications.
method Antithetic and Monte Carlo kernel estimators for partial rankings, variance reduction scheme.
result Improved antithetic kernel estimator with lower variance and better performance.
Geodesics found in deep linear networks.
problem Finding shortest paths in deep neural networks.
method Derived ODEs and explicit solutions for geodesics.
result Horizontal straight lines are geodesics in invariant manifold.
New geometric description of matrix manifolds avoiding equivalence classes.
problem Geometric description of matrix manifolds of fixed rank.
method Introducing a new geometric description of manifolds of matrices of fixed rank, avoiding equivalence classes.
result The matrix space Rnimesm is described as an analytic manifold equipped with a topology for which the matrix rank is a continuous map. New model for high rank matrix completion with online and batch methods.
problem Matrix completion for high rank matrices with latent structure.
method Kernel trick to map data into a high dimensional feature space, explicit parametrization of low dimensional subspace, online fitting procedure.
result Online method can handle streaming data and adapt to non-stationary latent structure.
The study examines how adding noise to neural networks improves reaching global optima.
problem Improving the optimization of deep neural networks.
method Theoretical analysis of noise's impact on the trajectories of gradient descent in multi-layer linear neural networks.
result Adding noise to a neural network increases the rank of the product of weight matrices, aiding in reaching a global optimum.
New method solves nonsmooth low-rank matrix optimization problems efficiently.
problem Nonsmooth and low-rank matrix optimization problems in statistics and machine learning.
method Low-rank Extragradient Method with warm-start initialization.
result The extragradient method converges to an optimal solution with rate O(1/t) and requires only two low-rank SVDs per iteration. The paper analyzes how low-rank layers in neural networks improve generalization.
problem Understanding how low-rank layers affect generalization in neural networks.
method Applying Maurer's chain rule for Gaussian complexity to analyze rank and spectral norm constraints.
result Deep networks with low-rank layers achieve better generalization than those with full-rank layers.
New algorithm tackles low-rank constraints in optimal transport problems.
problem Optimal transport problems with low-rank constraints.
method Explicit factorization of low-rank couplings as a product of sub-coupling factors linked by a common marginal.
result Stationary convergence of the algorithm proved.
New metrics improve landing algorithms for orthogonality constraints.
problem Optimizing landing algorithms with orthogonality constraints.
method Proposed a family of metrics over full-rank matrices to enhance landing algorithms.
result Natural extension of β-metric improves landing performance.
Optimizes neural network training by dynamically updating Tucker decomposition ranks.
problem Redundant parameters in neural network architectures.
method Geometry-aware training of factorized layers in tensor Tucker format.
result Optimal locally approximating the original dynamics without initial rank knowledge.
Gradient descent works well for large NNs due to convexity in a transformed space.
problem Why gradient descent works well in non-convex NN optimization.
method Introduced canonical space and disparity matrix to prove convexity.
result Gradient descent converges to global minimum in large NNs.
Signals are generally modeled as a superposition of exponential functions in spectroscopy of chemistry, biology and medical imaging. For fast data acquisition or other inevitable reasons, however, only a small amount of samples may be acquired and thus how to recover the full signal becomes an active research topic. Bu…
Minimal submanifolds in matrix spaces proven for specific ranks.
problem Minimal submanifolds in matrix spaces.
method Proving semialgebraic sets of matrices are minimal.
result Rectangular, skew-symmetric, and symmetric matrices with prescribed eigenvalues are minimal.
Rank-one measurements limit feasible sets for low-rank PSD matrices.
problem Feasibility of PSD matrices under rank-one measurements.
method Characterization of feasible sets for PSD matrices given rank-one projections.
result Radius of feasible sets determines singleton solution sets for low-rank matrices.
This paper tackles fitting multilevel low rank matrices by addressing three problems.
problem Fitting a given matrix by an MLR matrix in the Frobenius norm.
method Factor fitting, rank allocation, and hierarchical partitioning.
result The proposed methods can fit a given matrix by an MLR matrix in the Frobenius norm.
Decomposable-Net compresses neural networks without retraining for various sizes.
problem Performance degradation when changing model size after training.
method Decomposes weight matrices via SVD and adjusts ranks for different sizes.
result Maintains and improves performance across multiple model sizes.
New analysis shows how attention masks and LayerNorm prevent rank collapse in transformers.
problem Rank collapse in transformer models with increasing depth.
method General analysis of rank collapse under self-attention, considering attention masks and LayerNorm.
result Self-attention with LayerNorm can prevent rank collapse and maintain a rich set of equilibria.
New framework finds more efficient linear layers over structured matrices.
problem Efficient alternatives for dense linear layers in neural networks.
method Unified framework searching over all linear operators, developing a taxonomy based on computational and algebraic properties.
result BTT-MoE provides substantial compute-efficiency gains over dense layers and standard MoE.
Develops log-Euclidean Lie groups for SPD and correlation matrices.
problem Unifies various log-Euclidean constructions for SPD and correlation matrices.
method Theory and explicit isometries linking different log-Euclidean metrics.
result Explicit log-Euclidean metrics on SPD and correlation matrices.
A necessary condition for a connection in a vector bundle to be locally metric is for its curvature matrix, which consists of 2 forms, to be skew symmetric with respect to some local frame. In this paper we give a simple algorithm that can be used to decide when a matrix of 2 forms is equivalent to a skew symmetric…
Full-capacity uRNNs improve performance over restricted-capacity ones.
problem Vanishing and exploding gradient issues in recurrent neural networks.
method Optimized full-capacity unitary recurrence matrices over all unitary matrices.
result Significantly improved performance compared to LSTMs and restricted-capacity uRNNs.
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
problem Efficiently reducing the rank of non-negative matrices.
method Formulated rank reduction as a mean-field approximation using a log-linear model.
result Optimal solution for minimizing KL divergence can be computed in closed form.
New methods predict brain age from MEG/EEG without source modeling.
problem Predicting brain age from MEG/EEG data without source localization.
method Two Riemannian approaches to vectorize rank-reduced covariance matrices for regression.
result Data-driven Riemannian methods outperform sensor-space estimators and biophysics models.
New algorithms improve RPCA for large matrices with upper rank bounds.
problem Efficiently decompose large matrices into low-rank and sparse parts.
method Combine regularization and matrix multiplication approaches with upper rank bounds.
result Proposed algorithms are faster and more robust than existing methods.
New algorithm learns low-rank matrices with linear number of samples.
problem Learning low-rank matrices efficiently in latent-variable applications.
method Proposed algorithm that uses linear number of samples in high dimension.
result Learning kimesk, rank-r, matrices requires $Ω(rac{kr}{ε^2})$ samples. Low-rank MPPCA improves importance sampling in high dimensions.
problem Estimating full-rank GMM covariance matrices in high dimensions is numerically unstable.
method Use MPPCA mixtures as low-rank proposals for importance sampling in high-dimensional spaces.
result Consistent gains in sample efficiency and quality of failure distribution characterization.
This paper studies geodesics between covariance matrices of different ranks using the Bures-Wasserstein metric.
problem Geodesics between covariance matrices of varying ranks.
method Analyzes the Bures-Wasserstein distance on covariance matrices, completing previous work on geodesics and providing explicit formulas.
result The set of all minimizing geodesics between two covariance matrices is parametrized by a closed unit ball in R(k−r)imes(l−r).