Recently, path norm was proposed as a new capacity measure for neural networks with Rectified Linear Unit (ReLU) activation function, which takes the rescaling-invariant property of ReLU into account. It has been shown that the generalization error bound in terms of the path norm explains the empirical generalization b…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A toolkit for path-norms enhances neural network generalization bounds.
Deep networks with path norm regularization can approximate analytic functions.
PSiLON Net uses weight normalization and 1-path-norm regularization for efficient learning and sparsity.
Complexity measures for neural nets with general activations using path-based norms.
New method for efficient proximal mapping of 1-path-norm in shallow networks.
A new method detects anomalies in multivariate streams without unit dependence.
Optimal a priori estimates are derived for the population risk, also known as the generalization error, of a regularized residual network model. An important part of the regularized model is the usage of a new path norm, called the weighted path norm, as the regularization term. The weighted path norm treats the skip c…
Convexity proven for sums of angles of unitary paths.
We introduce here a natural functional associated to any : \emph{spectral length functional}, on the space of "generalized paths" in , closely related to both the Hofer length functional and spectral invariants and establish some of its properties. This functional is smooth on its…
Diagonal linear networks converge to lasso regularization path during training.
Algorithm approximates regularization path for deep neural networks efficiently.
Efficient algorithms for clustered Lasso and OSCAR reduce computational costs.
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with …
We use the criteria of Lalonde and McDuff to determine a new class of examples of length minimizing paths in the group . For a compact symplectic manifold of dimension two or four, we show that a path in , generated by an autonomous Hamiltonian and starting at the identity, which induces no non-cons…
This paper analyzes shallow ReLU networks in L^p and Sobolev spaces, focusing on approximation and generalization.
Study shows how feature weighting affects neural network regularization.
The -1 norm based optimization is widely used in signal processing, especially in recent compressed sensing theory. This paper studies the solution path of the -1 norm penalized least-square problem, whose constrained form is known as Least Absolute Shrinkage and Selection Operator (LASSO). A solution path …
This paper concerns model reduction of dynamical systems using the nuclear norm of the Hankel matrix to make a trade-off between model fit and model complexity. This results in a convex optimization problem where this trade-off is determined by one crucial design parameter. The main contribution is a methodology to app…
Lie groups with bi-invariant distance are products of abelian and compact groups.
Global approximation for piecewise linear paths via signatures.
This paper presents a normalization mechanism called Instance-Level Meta Normalization (ILM~Norm) to address a learning-to-normalize problem. ILM~Norm learns to predict the normalization parameters via both the feature feed-forward and the gradient back-propagation paths. ILM~Norm provides a meta normalization mechanis…
This paper studies neural networks with bounded norms to avoid the curse of dimensionality.
A new framework uses matrix flows to unify frequentist and Bayesian approaches for sparse GGMs.
Max-convolution is an important problem closely resembling standard convolution; as such, max-convolution occurs frequently across many fields. Here we extend the method with fastest known worst-case runtime, which can be applied to nonnegative vectors by numerically approximating the Chebyshev norm $\| \cdot \|_\infty…
Geodesics of contactomorphisms on a specific manifold are characterized by Hamiltonian functions.
Quaternionic frames' admissibility and homotopy proven.
We use convex relaxation techniques to provide a sequence of solutions to the matrix completion problem. Using the nuclear norm as a regularizer, we provide simple and very efficient algorithms for minimizing the reconstruction error subject to a bound on the nuclear norm. Our algorithm iteratively replaces the missing…
New algorithm speeds up path computation for optimal models.
Develops a new solver for path-dependent PDEs using signature kernels.
We develop a normative framework for hierarchical model-based policy optimization based on applying second-order methods in the space of all possible state-action paths. The resulting natural path gradient performs policy updates in a manner which is sensitive to the long-range correlational structure of the induced st…
We consider the generic regularized optimization problem . Efron, Hastie, Johnstone and Tibshirani [Ann. Statist. 32 (2004) 407--499] have shown that for the LASSO--that is, if is squared error loss and is the norm of --the opti…
Left invariant metrics induced by the p-norms of the trace in the matrix algebra are studied on the general lineal group. By means of the Euler-Lagrange equations, existence and uniqueness of extremal paths for the length functional are established, and regularity properties of these extremal paths are obtained. Minimi…
The relaxed maximum entropy problem is concerned with finding a probability distribution on a finite set that minimizes the relative entropy to a given prior distribution, while satisfying relaxed max-norm constraints with respect to a third observed multinomial distribution. We study the entire relaxation path for thi…
New Banach spaces for ReLU networks enable better function approximation and gradient dynamics analysis.
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an -path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regula…
We study adaptive regret bounds in terms of the variation of the losses (the so-called path-length bounds) for both multi-armed bandit and more generally linear bandit. We first show that the seemingly suboptimal path-length bound of (Wei and Luo, 2018) is in fact not improvable for adaptive adversary. Despite this neg…
Existence of Q-processes for Brownian motion on hyperbolic spaces with Poissonian potentials shown.
We study the complexity of the entire regularization path for least squares regression with 1-norm penalty, known as the Lasso. Every regression parameter in the Lasso changes linearly as a function of the regularization value. The number of changes is regarded as the Lasso's complexity. Experimental results using exac…
In this paper we first show that the necessary condition introduced in our previous paper is also a sufficient condition for a path to be a geodesic in the group $\Ham^c(M)$ of compactly supported Hamiltonian symplectomorphisms. This applies with no restriction on . We then discuss conditions which guarantee that su…
Unified representation for tree ensembles indexed by nodes
Consider the group $\Ham^c(M)$ of compactly supported Hamiltonian symplectomorphisms of the symplectic manifold $(M,\om)$ with the Hofer -norm. A path in $\Ham^c(M)$ will be called a geodesic if all sufficiently short pieces of it are local minima for the Hofer length functional $\Ll$. In this paper, we giv…
Path regularization reveals convex optimization in deep ReLU networks.
Classical results on the statistical complexity of linear models have commonly identified the norm of the weights as a fundamental capacity measure. Generalizations of this measure to the setting of deep networks have been varied, though a frequently identified quantity is the product of weight norms of each la…
The paper proves parabolic gap theorems for Yang-Mills energy.
GNMR controls runtime stability in low-precision language model training.
This paper certifies cluster assignments from sum-of-norms clustering algorithms.
Let stand for the path connected identity component of the group of all compactly supported homeomorphisms of a manifold . It is shown that is perfect and simple under mild assumptions on . Next, conjugation-invariant norms on $\H_c(M)$ are considered and the boundedness of $\m…