The spectral -support norm enjoys good estimation properties in low rank matrix learning problems, empirically outperforming the trace norm. Its unit ball is the convex hull of rank matrices with unit Frobenius norm. In this paper we generalize the norm to the spectral -support norm, whose additional para…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Exact spectral norm regularization improves neural network generalization.
We study a regularizer which is defined as a parameterized infimum of quadratics, and which we call the box-norm. We show that the k-support norm, a regularizer proposed by [Argyriou et al, 2012] for sparse vector prediction problems, belongs to this family, and the box-norm can be generated as a perturbation of the fo…
We investigate the generalizability of deep learning based on the sensitivity to input perturbation. We hypothesize that the high sensitivity to the perturbation of data degrades the performance on it. To reduce the sensitivity to perturbation, we propose a simple and effective regularization method, referred to as spe…
Equivalence of norms on manifolds with curvature bounds established.
We introduce here a natural functional associated to any : \emph{spectral length functional}, on the space of "generalized paths" in , closely related to both the Hofer length functional and spectral invariants and establish some of its properties. This functional is smooth on its…
In deep neural networks, the spectral norm of the Jacobian of a layer bounds the factor by which the norm of a signal changes during forward/backward propagation. Spectral norm regularizations have been shown to improve generalization, robustness and optimization of deep learning methods. Existing methods to compute th…
The study connects norms and filtrations on section rings of projective manifolds.
Muon dynamics study uses spectral Wasserstein flow for optimization stability.
Spectral regularization simplifies sequence models by focusing on grammatical simplicity.
We show that the spectral norm of a random tensor (or higher-order array) scales as under some sub-Gaussian assumption on the entries. The proof is based on a covering number argument. Since the spectral norm is dual to the tensor…
Muon optimizer improves deep learning with spectral norm constraints.
Study on neural networks' sample complexity with one hidden layer.
Improved singular value approximation for convolutional layers.
New bounds adaptively control spectral complexity of trained Transformers.
Extended Gauss-Markov theorem for linear estimation with bounded bias.
Study spectral distribution of twisted Laplacian on high genus hyperbolic surfaces.
The study proves properties of spectral selectors for contact manifolds and applies them to contact big fibers and geodesics.
The paper improves tensor completion bounds using spectral gap.
Optimal estimates for spectral projection norms on compact manifolds.
In this dissertation we propose alternative analysis of distributed stochastic gradient descent (SGD) algorithms that rely on spectral properties of the data covariance. As a consequence we can relate questions pertaining to speedups and convergence rates for distributed SGD to the data distribution instead of the regu…
The study improves norms of spectral projectors on specific surfaces.
Proves spectral gap bounds for Teichmüller geodesics on flat surfaces.
Algorithm estimates covariance from noisy data efficiently.
In this paper, we consider the Tensor Robust Principal Component Analysis (TRPCA) problem, which aims to exactly recover the low-rank and sparse components from their sum. Our model is based on the recently proposed tensor-tensor product (or t-product). Induced by the t-product, we first rigorously deduce the tensor sp…
Recent research in off-the-grid compressed sensing (CS) has demonstrated that, under certain conditions, one can successfully recover a spectrally sparse signal from a few time-domain samples even though the dictionary is continuous. In particular, atomic norm minimization was proposed in \cite{tang2012csotg} to recove…
Paper provides a performance guarantee for spectral clustering.
This paper proves a generalization bound for complex-valued neural networks scaling with spectral complexity.
Study precise sample covariance error for Gaussian centered data.
Study extends neural network approximation to time-varying PDEs using Fourier-Lebesgue spaces.
Estimates spectral projections restricted to uniformly embedded submanifolds.
Defines spectral selectors on lens spaces for contactomorphisms.
Study of unitary and groupoid orbits of normal operators, focusing on manifold structures and spectral conditions.
In this paper, we consider low rank matrix estimation using either matrix-version Dantzig Selector or matrix-version LASSO estimator . We consider sub-Gaussian measurements, , the measurements have sub-Gaussian entries. Suppose $\textrm…
The extraction of clusters from a dataset which includes multiple clusters and a significant background component is a non-trivial task of practical importance. In image analysis this manifests for example in anomaly detection and target detection. The traditional spectral clustering algorithm, which relies on the lead…
For graphs generated from stochastic blockmodels, adjacency spectral embedding is asymptotically consistent. Further, adjacency spectral embedding composed with universally consistent classifiers is universally consistent to achieve the Bayes error. However when the graph contains private or sensitive information, trea…
The -support norm is a regularizer which has been successfully applied to sparse vector prediction problems. We show that it belongs to a general class of norms which can be formulated as a parameterized infimum over quadratics. We further extend the -support norm to matrices, and we observe that it is a special …
We establish a theoretical link between adversarial training and operator norm regularization for deep neural networks. Specifically, we prove that -norm constrained projected gradient ascent based adversarial training with an -norm loss on the logits of clean and perturbed inputs is equivalent to data-…
A new method speeds up spectral normalization for neural nets.
Paper provides robustness bounds for GNNs against adversarial attacks.
Paper finds exact Hessian sharpness in deep matrix factorization.
We apply the integral formula of volumes to the family of graded linear series constructed from any test configuration. This solves the conjecture raised by Witt--Nyström so that the sequence of spectral measures for the induced -action on the central fiber converges to the canonical Duistermatt--Heckman …
Correlation matrices play a key role in many multivariate methods (e.g., graphical model estimation and factor analysis). The current state-of-the-art in estimating large correlation matrices focuses on the use of Pearson's sample correlation matrix. Although Pearson's sample correlation matrix enjoys various good prop…
We propose a new point of view for regularizing deep neural networks by using the norm of a reproducing kernel Hilbert space (RKHS). Even though this norm cannot be computed, it admits upper and lower approximations leading to various practical strategies. Specifically, this perspective (i) provides a common umbrella f…
Let be a compact, oriented 3-manifold with a contact form and a metric . Suppose that is a principal bundle with structure group such that is the principal SO(3) bundle of orthonormal frames for . A unitary connection on the Hermitian line bundle $…
New algorithm for multiway spectral clustering on Grassmann manifolds.
Proves singular support of sheaves is γ-coisotropic, with implications for symplectic homeomorphisms.
This paper is concerned about sparse, continuous frequency estimation in line spectral estimation, and focused on developing gridless sparse methods which overcome grid mismatches and correspond to limiting scenarios of existing grid-based approaches, e.g., optimization and SPICE, with an infinitely dense grid…