A novel metric and framework for evaluating gradient norm equality in deep neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A popular heuristic for improved performance in Generative adversarial networks (GANs) is to use some form of gradient penalty on the discriminator. This gradient penalty was originally motivated by a Wasserstein distance formulation. However, the use of gradient penalty in other GAN formulations is not well motivated.…
It has been noted in existing literature that over-parameterization in ReLU networks generally improves performance. While there could be several factors involved behind this, we prove some desirable theoretical properties at initialization which may be enjoyed by ReLU networks. Specifically, it is known that He initia…
Muon replaces matrix gradient with polar factor, optimizing flat spectrum updates
The article derives some novel independence measures and contrast functions for Blind Source Separation (BSS) application. For the order differentiable multivariate functions with equal hyper-volumes (region bounded by hyper-surfaces) and with a constraint of bounded support for , it proves that equality …
This work analyzes the maximum-margin bias in quasi-homogeneous neural networks.
In this paper, we prove that a noncompact complete hypersurface with finite weighted volume, weighted mean curvature vector bounded in norm, and isometrically immersed in a complete weighted manifold is proper. In addition, we obtain an estimate for -stability index of a constant weighted mean curvature hypersurface…
There are different problems for resolution of complex LC-MS or GC-MS data, such as the existence of embedded chromatographic peaks, continuum background and overlapping in mass channels for different components. These problems cause rotational ambiguity in recovered profiles calculated using multivariate curve resolut…
We obtain a Chern-Osserman type equality of a complete properly immersed surface in Euclidean space, provided the L^2-norm of the second fundamental form is finite. Also, by using a monotonicity formula, we prove that if the L^2-norm of mean curvature of a noncompact surface is finite, then it has at least quadratic ar…
Deep learning algorithms have increasingly been shown to lack robustness to simple adversarial examples (AdvX). An equally troubling observation is that these adversarial examples transfer between different architectures trained on different datasets. We investigate the transferability of adversarial examples between m…
New bound relaxes uniform gradient norm assumptions for PAC-Bayesian bounds.
Paper characterizes nc-rank using gradient flow on symmetric space.
Gradient descent on Hadamard manifolds converges to boundary points, solving optimization problems.
We relate generalized Lebesgue decompositions of measures in terms of curve fragments (Alberti representations) and Weaver derivations. This correspondence leads to a geometric characterization of the local norm on the Weaver cotangent bundle of a metric measure space : the local norm of a form sees how fas…
We prove certain optimal systolic inequalities for a closed Riemannian manifold (X,g), depending on a pair of parameters, n and b. Here n is the dimension of X, while b is its first Betti number. The proof of the inequalities involves constructing Abel-Jacobi maps from X to its Jacobi torus T^b, which are area-decreasi…
Two Dehn surgeries on a knot are called cosmetic if they yield homeomorphic manifolds. For a null-homologous knot with certain conditions on the Thurston norm of the ambient manifold, if the knot admits cosmetic surgeries, then the surgery coefficients are equal up to sign.
AdaGrad-Norm achieves optimal convergence rates for non-convex objectives without tuning.
This technical report describes an efficient technique for computing the norm of the gradient of the loss function for a neural network with respect to its parameters. This gradient norm can be computed efficiently for every example.
Equivalent tests for SGD batch size selection found.
The paper proves stability and convergence of minimal networks under curvature motion.
New algorithms adapt to both gradient norms and comparator norms in online learning.
Adam's hyperparameters implicitly regularize solutions, penalizing or impeding loss gradients' norms.
We prove that the norm version of the adaptive stochastic gradient method (AdaGrad-Norm) achieves a linear convergence rate for a subset of either strongly convex functions or non-convex functions that satisfy the Polyak Lojasiewicz (PL) inequality. The paper introduces the notion of Restricted Uniform Inequality of Gr…
Studied how SGD's stability regularization affects generalization in neural networks.
A method predicts GNS of transformer layers using normalization layer norms.
Improved greedy 2-coordinate updates for optimization problems with constraints.
New SPS variant improves non-smooth optimization without small gradients.
The study determines -Thurston norms in Sol manifolds and embeds non-orientable surfaces.
FPGA-based multi-layer equalizer adapts to changing channels.
Gradient flow on softmax attention minimizes nuclear norm of weight matrices.
Training neural networks under a strict Lipschitz constraint is useful for provable adversarial robustness, generalization bounds, interpretable gradients, and Wasserstein distance estimation. By the composition property of Lipschitz functions, it suffices to ensure that each individual affine transformation or nonline…
We provide a parametric construction in terms of minimal surfaces of the Euclidean submanifolds of codimension two and arbitrary dimension that attain equality in an inequality due to De Smet, Dillen, Verstraelen and Vrancken. The latter involves the scalar curvature, the norm of the normal curvature tensor and the len…
Generalizes Reilly inequality to varifolds and analyzes equality cases.
We address some theoretical guarantees for Schatten- quasi-norm minimization () in recovering low-rank matrices from compressed linear measurements. Firstly, using null space properties of the measurement operator, we provide a sufficient condition for exact recovery of low-rank matrices. This condition…
This work improves OOD detection using deep generative models by approximating Fisher information metrics.
Dropout and its extensions (eg. DropBlock and DropConnect) are popular heuristics for training neural networks, which have been shown to improve generalization performance in practice. However, a theoretical understanding of their optimization and regularization properties remains elusive. Recent work shows that in the…
We model how Lipschitz continuity changes during neural network training.
Noise injection before gradient steps helps in regularization for neural networks.
We study a hybrid conditional gradient - smoothing algorithm (HCGS) for solving composite convex optimization problems which contain several terms over a bounded set. Examples of these include regularization problems with several norms as penalties and a norm constraint. HCGS extends conditional gradient methods to cas…
Gaussian Markov random fields (GMRFs) are useful in a broad range of applications. In this paper we tackle the problem of learning a sparse GMRF in a high-dimensional space. Our approach uses the l1-norm as a regularization on the inverse covariance matrix. We utilize a novel projected gradient method, which is faster …
RTC-GTNLN model recovers traffic data from missing values and noise.
New method uses causal thinking to make AI fairer decisions.
For a 3-manifold M, McMullen derived from the Alexander polynomial of M a norm on H^1(M, R) called the Alexander norm. He showed that the Thurston norm on H^1(M, R), which measures the complexity of a dual surface, is an upper bound for the Alexander norm. He asked if these two norms were equal on all of H^1(M,R) when …
Gradient flow with weight decay shows grokking effect in deep learning.
Unified framework approximates gradient descent's implicit bias in high dimensions.
Let M be a bounded open plane domain. Let f be a continuous function on the closure of M, 3-times continuously differentiable in M, which vanish on the boundary. Polterovich and Sodin proved that the values of f cannot exceed the norm of the hessian of f, averaged over the entire domain M. In this paper we study the eq…
For a compact, orientable, irreducible 3-manifold with toroidal boundary that is not the product of a torus and an interval or a cable space, each boundary torus has a finite set of slopes such that, if avoided, the Thurston norm of a Dehn filling behaves predictably. More precisely, for all but finitely many slopes, t…
We study the iteration complexity of stochastic gradient descent (SGD) for minimizing the gradient norm of smooth, possibly nonconvex functions. We provide several results, implying that the upper bound of Ghadimi and Lan~\cite{ghadimi2013stochastic} (for making the average gradient norm less than…