The Schatten quasi-norm can be used to bridge the gap between the nuclear norm and rank function, and is the tighter approximation to matrix rank. However, most existing Schatten quasi-norm minimization (SQNM) algorithms, as well as for nuclear norm minimization, are too slow or even impractical for large-scale problem…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper develops a new class of nonconvex regularizers for low-rank matrix recovery. Many regularizers are motivated as convex relaxations of the matrix rank function. Our new factor group-sparse regularizers are motivated as a relaxation of the number of nonzero columns in a factorization of the matrix. These nonco…
New proof shows norms can't explain deep learning's implicit regularization.
Paper tackles low-rank matrix recovery with column -norm regularization.
New method for factor analysis using nuclear and norms.
This work shows how penalising bias terms in norm regularisation leads to sparse solutions.
This paper studies optimal approximation factors in misspecified off-policy RL, identifying key factors under various settings.
The Schatten- norm () has been widely used to replace the nuclear norm for better approximating the rank function. However, existing methods are either 1) not scalable for large scale problems due to relying on singular value decomposition (SVD) in every iteration, or 2) specific to some values, e.g., $1/…
Based on a new atomic norm, we propose a new convex formulation for sparse matrix factorization problems in which the number of nonzero elements of the factors is assumed fixed and known. The formulation counts sparse PCA with multiple factors, subspace clustering and low-rank sparse bilinear regression as potential ap…
This paper is concerned with the squared F(robenius)-norm regularized factorization form for noisy low-rank matrix recovery problems. Under a suitable assumption on the restricted condition number of the Hessian for the loss function, we derive an error bound to the true matrix for the non-strict critical points with r…
Paper finds exact Hessian sharpness in deep matrix factorization.
In this note, we derive concentration inequalities for random vectors with subGaussian norm (a generalization of both subGaussian random vectors and norm bounded random vectors), which are tight up to logarithmic factors.
New ONMF model minimizes KL divergence for better sparse data modeling.
Study shows interpolating predictor's risk is optimal in low-dimensional factor regression models.
This paper is concerned with the factorization form of the rank regularized loss minimization problem. To cater for the scenario in which only a coarse estimation is available for the rank of the true matrix, an -norm regularized term is added to the factored loss function to reduce the rank adaptively; and…
Unified framework for estimating high-dimensional conditional factor models.
DoRA improves adaptation efficiency for large models by factoring norms and fusing kernels.
We make an estimation of the value of the Gromov norm of the Cartesian product of two surfaces. Our method uses a connection between these norms and the minimal size of triangulations of the products of two polygons. This allows us to prove that the Gromov norm of this product is between 32 and 52 when both factors hav…
We study the projected gradient descent method on low-rank matrix problems with a strongly convex objective. We use the Burer-Monteiro factorization approach to implicitly enforce low-rankness; such factorization introduces non-convexity in the objective. We focus on constraint sets that include both positive semi-defi…
New algorithms improve approximation of matrix norms, with applications in statistics and machine learning.
A new model reduces noise and speeds up subspace segmentation.
New model reduces matrix factorization bias, yielding truly low-rank solutions.
Gradient flow with infinitesimal initialization converges to Greedy Low-Rank Learning for matrix factorization.
The heavy-tailed distributions of corrupted outliers and singular values of all channels in low-level vision have proven effective priors for many applications such as background modeling, photometric stereo and image alignment. And they can be well modeled by a hyper-Laplacian. However, the use of such distributions g…
We study implicit regularization when optimizing an underdetermined quadratic objective over a matrix with gradient descent on a factorization of . We conjecture and provide empirical and theoretical evidence that with small enough step sizes and initialization close enough to the origin, gradient descent on a f…
Regularization for matrix factorization (MF) and approximation problems has been carried out in many different ways. Due to its popularity in deep learning, dropout has been applied also for this class of problems. Despite its solid empirical performance, the theoretical properties of dropout as a regularizer remain qu…
Dropout is a simple yet effective algorithm for regularizing neural networks by randomly dropping out units through Bernoulli multiplicative noise, and for some restricted problem classes, such as linear or logistic regression, several theoretical studies have demonstrated the equivalence between dropout and a fully de…
Improved sample complexity for ReLU networks with norm constraints.
Paper analyzes convergence of PAM method for low-rank factorization models.
We model how Lipschitz continuity changes during neural network training.
Recovering low-rank and sparse matrices from incomplete or corrupted observations is an important problem in machine learning, statistics, bioinformatics, computer vision, as well as signal and image processing. In theory, this problem can be solved by the natural convex joint/mixed relaxations (i.e., l_{1}-norm and tr…
This paper tackles fitting multilevel low rank matrices by addressing three problems.
In this note we determine all possible dominations between different products of manifolds, when none of the factors of the codomain is dominated by products. As a consequence, we determine the finiteness of every product-associated functorial semi-norm on the fundamental classes of the aforementioned products. These r…
New bounds improve deep learning performance efficiently.
We propose an algorithm for the non-negative factorization of an occurrence tensor built from heterogeneous networks. We use l0 norm to model sparse errors over discrete values (occurrences), and use decomposed factors to model the embedded groups of nodes. An efficient splitting method is developed to optimize the non…
The Schatten quasi-norm was introduced to bridge the gap between the trace norm and rank function. However, existing algorithms are too slow or even impractical for large-scale problems. Motivated by the equivalence relation between the trace norm and its bilinear spectral penalty, we define two tractable Schatten norm…
New bounds adaptively control spectral complexity of trained Transformers.
The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.
Unified framework for coupled tensor completion improves recovery accuracy.
We study the adaptive estimation of copula correlation matrix for the semi-parametric elliptical copula model. In this context, the correlations are connected to Kendall's tau through a sine function transformation. Hence, a natural estimate for is the plug-in estimator with Kendall's tau statistic. We …
The study analyzes robustness of estimators in linear models with adversarial errors.
Max-norm regularizer has been extensively studied in the last decade as it promotes an effective low-rank estimation for the underlying data. However, such max-norm regularized problems are typically formulated and solved in a batch manner, which prevents it from processing big data due to possible memory budget. In th…
Recently, locality sensitive hashing (LSH) was shown to be effective for MIPS and several algorithms including -ALSH, Sign-ALSH and Simple-LSH have been proposed. In this paper, we introduce the norm-range partition technique, which partitions the original dataset into sub-datasets containing items with similar 2-…
Paper shows no spurious local minima in a specific matrix factorization problem.
Gradient descent promotes low-rank solutions in tensor completion.
New bounds show current methods overestimate system parameter errors.
The paper tackles tensor factorization and completion from noisy data.
The paper studies the metric and algebraic structures on section rings of projective manifolds.