We introduce a new family of matrix norms, the "local max" norms, generalizing existing methods such as the max norm, the trace norm (nuclear norm), and the weighted or smoothed weighted trace norms, which have been extensively used in the literature as regularizers for matrix reconstruction problems. We show that this…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recently theoretical guarantees have been obtained for matrix completion in the non-uniform sampling regime. In particular, if the sampling distribution aligns with the underlying matrix's leverage scores, then with high probability nuclear norm minimization will exactly recover the low rank matrix. In this article, we…
Sharp inequality on Siegel domain involving weighted norms and sub-Laplacian.
In recent studies, several asymptotic upper bounds on generalization errors on deep neural networks (DNNs) are theoretically derived. These bounds are functions of several norms of weights of the DNNs, such as the Frobenius and spectral norms, and they are computed for weights grouped according to either input and outp…
Gradient flow on softmax attention minimizes nuclear norm of weight matrices.
This work shows how penalising bias terms in norm regularisation leads to sparse solutions.
PSiLON Net uses weight normalization and 1-path-norm regularization for efficient learning and sparsity.
We provide rigorous guarantees on learning with the weighted trace-norm under arbitrary sampling distributions. We show that the standard weighted trace-norm might fail when the sampling distribution is not a product distribution (i.e. when row and column indexes are not selected independently), present a corrected var…
This is the fourth article of our series. Here, we study weighted norm inequalities for the Riesz transform of the Laplace-Beltrami operator on Riemannian manifolds and of subelliptic sum of squares on Lie groups, under the doubling volume property and Gaussian upper bounds.
Characterizes inductive bias in multi-channel linear CNNs with bounded weight norm.
Any Sasakian structure can be closely mimicked by embeddings into weighted spheres.
To recover a sparse signal from an underdetermined system, we often solve a constrained L1-norm minimization problem. In many cases, the signal sparsity and the recovery performance can be further improved by replacing the L1 norm with a "weighted" L1 norm. Without any prior information about nonzero elements of the si…
New method stabilizes FQE by reweighting Bellman targets.
Over the past few years, Batch-Normalization has been commonly used in deep networks, allowing faster training and high performance for a wide variety of applications. However, the reasons behind its merits remained unanswered, with several shortcomings that hindered its use for certain tasks. In this work, we present …
Theoretical justification for deep networks' performance with regularization techniques.
Gradient flow with weight decay shows grokking effect in deep learning.
Proposes a new regression method using -norms for non-Gaussian noise.
In recent years, the nuclear norm minimization (NNM) problem has been attracting much attention in computer vision and machine learning. The NNM problem is capitalized on its convexity and it can be solved efficiently. The standard nuclear norm regularizes all singular values equally, which is however not flexible enou…
Optimal a priori estimates are derived for the population risk, also known as the generalization error, of a regularized residual network model. An important part of the regularized model is the usage of a new path norm, called the weighted path norm, as the regularization term. The weighted path norm treats the skip c…
The paper predicts survival functions using random survival trees and concordance maximization.
Deep networks with path norm regularization can approximate analytic functions.
A new method to improve deep neural networks using weight rescaling.
The paper proposes a novel MKL approach for OCC using -norm constraints.
Deep neural networks (DNNs) have become increasingly important due to their excellent empirical performance on a wide range of problems. However, regularization is generally achieved by indirect means, largely due to the complex set of functions defined by a network and the difficulty in measuring function complexity. …
Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have been widely used in modern machine learning. Here, we study the weight normalization (WN) method [Salimans and Kingma, 2016] and a variant call…
New regularizer improves neural network robustness and generalization.
The paper analyzes methods for estimating linear functionals from observational data, proving upper bounds and showing optimal procedures.
A key element of understanding the efficacy of overparameterized neural networks is characterizing how they represent functions as the number of weights in the network approaches infinity. In this paper, we characterize the norm required to realize a function as a single hidden-lay…
This work is substituted by the paper in arXiv:2011.14066. Stochastic gradient descent is the de facto algorithm for training deep neural networks (DNNs). Despite its popularity, it still requires fine tuning in order to achieve its best performance. This has led to the development of adaptive methods, that claim autom…
In this short report, we discuss how coordinate-wise descent algorithms can be used to solve minimum variance portfolio (MVP) problems in which the portfolio weights are constrained by norms, where . A portfolio which weights are regularised by such norms is called a sparse portfolio (Brodie et …
RVFL networks can efficiently approximate Lipschitz functions in L∞ norm.
AdamW optimizes a constrained loss with norm constraint.
Multiple kernel learning (MKL), structured sparsity, and multi-task learning have recently received considerable attention. In this paper, we show how different MKL algorithms can be understood as applications of either regularization on the kernel weights or block-norm-based regularization, which is more common in str…
Study Gaussian approximation for deep neural networks with random weights.
A recent analysis of a model of iterative neural network in Hilbert spaces established fundamental properties of such networks, such as existence of the fixed points sets, convergence analysis, and Lipschitz continuity. Building on these results, we show that under a single mild condition on the weights of the network,…
New findings on depth vs. width in neural networks, showing depth can improve learnability.
This paper presents a general framework for norm-based capacity control for weight normalized deep neural networks. We establish the upper bound on the Rademacher complexities of this family. With an normalization where , and , we discuss properties of a width-independent ca…
SWRLDA improves LDA for multi-class classification with edge classes.
Classical results on the statistical complexity of linear models have commonly identified the norm of the weights as a fundamental capacity measure. Generalizations of this measure to the setting of deep networks have been varied, though a frequently identified quantity is the product of weight norms of each la…
Muon optimizer improves deep learning with spectral norm constraints.
We refine and generalize several interpolation inequalities bounding the norm of a probability density with respect to the reference measure by its Sobolev norm and the Kantorovich distance to on a smooth weighted Riemannian manifold satisfying condition.
Study shows how networks converge to minimum norm solutions with regularization.
New method approximates complex kernel norms with random features, making learning tractable.
In this paper, we propose a novel linear discriminant analysis criterion via the Bhattacharyya error bound estimation based on a novel L1-norm (L1BLDA) and L2-norm (L2BLDA). Both L1BLDA and L2BLDA maximize the between-class scatters which are measured by the weighted pairwise distances of class means and meanwhile mini…
Regularization can induce grokking in neural networks, improving generalization.
Kähler information manifolds for signal filters in weighted Hardy spaces are explored.
We investigate the generalizability of deep learning based on the sensitivity to input perturbation. We hypothesize that the high sensitivity to the perturbation of data degrades the performance on it. To reduce the sensitivity to perturbation, we propose a simple and effective regularization method, referred to as spe…
New theory maps neural network weights to optimize faster and scale.