PSiLON Net uses L1 weight normalization and 1-path-norm regularization for efficient learning and sparsity.
problem Efficient learning and sparsity in neural networks with limited data.
method PSiLON Net employs L1 weight normalization and 1-path-norm regularization to simplify the 1-path-norm and achieve efficient learning and near-sparse parameters. result PSiLON Net achieves reliable optimization and strong performance in the small data regime.
Deep networks with path norm regularization can approximate analytic functions.
problem Approximating analytic functions with neural networks.
method Path norm regularized deep networks with activation function.
result Deep networks can approximate analytic functions with logarithmic dependence on approximation error.
Optimal estimates derived for residual networks' generalization error.
problem Estimating the generalization error of residual networks.
method Derives optimal a priori estimates using a weighted path norm.
result Optimal error estimates are comparable to Monte Carlo error rates.
Study shows how feature weighting affects neural network regularization.
problem Understanding how feature weighting influences neural network regularization.
method Derived equivalence paths connecting different weighting matrices and ridge regularization levels.
result Ridge estimators trained on weighted features are asymptotically equivalent when evaluated against test vectors.
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with …
New capacity measure for deep ReLU networks derived from weight norms.
problem Identifying a suitable capacity measure for deep ReLU networks.
method Generalization of a recently proposed sampling argument to demonstrate the existence of sparse approximants of positive homogeneous networks.
result Bounding generalization error in multi-class classification using covering number bounds.
Global approximation for piecewise linear paths via signatures.
problem Global approximation theorems for piecewise linear paths.
method Using signatures of piecewise linear paths and their density in Lp-norms. result Linear functionals of signatures are dense in Lp-norms under an integrability condition. Diagonal linear networks converge to lasso regularization path during training.
problem Understanding the regularization behavior of diagonal linear networks.
method Analyzing the training trajectory of diagonal linear networks and comparing it to the lasso regularization path.
result The training trajectory of diagonal linear networks is closely related to the lasso regularization path.
Recently, path norm was proposed as a new capacity measure for neural networks with Rectified Linear Unit (ReLU) activation function, which takes the rescaling-invariant property of ReLU into account. It has been shown that the generalization error bound in terms of the path norm explains the empirical generalization b…
A toolkit for path-norms enhances neural network generalization bounds.
problem Establishing generalization bounds for modern neural networks.
method Introducing a comprehensive toolkit for path-norms in ReLU networks with various operations.
result Established generalization bounds for modern neural networks that are the most widely applicable and recover/beat the sharpest known bounds.
Algorithm approximates regularization path for deep neural networks efficiently.
problem Computing the regularization path for high-dimensional deep neural networks.
method Multiobjective continuation method for non-smooth objectives.
result Approximation of the entire Pareto front for regularization path.
To recover a sparse signal from an underdetermined system, we often solve a constrained L1-norm minimization problem. In many cases, the signal sparsity and the recovery performance can be further improved by replacing the L1 norm with a "weighted" L1 norm. Without any prior information about nonzero elements of the si…
Complexity measures for neural nets with general activations using path-based norms.
problem Control complexity of neural networks with arbitrary activation functions.
method Approximate general activations with ReLU networks and derive path-based norms for complexity control.
result Preliminary analyses of function spaces and regularized estimators.
New method for efficient proximal mapping of 1-path-norm in shallow networks.
problem Efficiently handling the 1-path-norm of shallow neural networks.
method Closed-form proximal operator for efficient computation and upper bound on Lipschitz constant.
result Proximal mapping allows robust training against adversarial perturbations.
This paper certifies cluster assignments from sum-of-norms clustering algorithms.
problem Certifying the correct cluster assignments from approximate solutions of sum-of-norms clustering.
method Presented a clustering test that identifies and certifies the correct cluster assignment from an approximate solution.
result The correct cluster assignment is guaranteed to be certified by a primal-dual path following algorithm after sufficient iterations.
A new method detects anomalies in multivariate streams without unit dependence.
problem Detect anomalies in multivariate streams without unit dependence.
method Proposes SigMahaKNN combining variance norm and path signature.
result SigMahaKNN detects anomalies better than existing methods.
The paper calculates sensitivities for financial derivatives using path weighting methods.
problem Computing sensitivities for path-dependent financial derivatives with high variance and degeneracy issues.
method Proposes explicit path weighting formula, variance reduction adjustment, and covariance inflation technique.
result Effective methods to address high variance and degeneracy in sensitivities computation.
Paper finds conditions for different norms to produce same billiard paths.
problem Conditions for different norms to define the same billiard reflection law.
method Extending previous works by Milena Radnović and Serge Tabachnikov, the paper establishes conditions for two different non-symmetric norms to define the same billiard reflection law.
result Conditions for two different norms to define the same billiard reflection law.
Convexity proven for sums of angles of unitary paths.
problem Proving convexity of sums of eigenvalues of unitary matrices.
method Analyzing paths of unitary matrices and their angles, using operator norms.
result Sum of first m angles of unitary path is convex.
Power weighted shortest paths improve clustering of high-dimensional data.
problem Clustering high-dimensional Euclidean data with disjoint low-dimensional manifolds.
method Use of power weighted shortest path distance functions and a fast algorithm.
result Higher clustering accuracy achieved through power weighted shortest paths.
Unified representation for tree ensembles indexed by nodes
problem Unifying geometric object for tree ensembles indexed by nodes
method KPP indexes feature map by nodes, weighted by path metric
result Unified non-diagonal Gram for prediction, additive attribution, robust radius, and risk bounds
We introduce here a natural functional associated to any b∈QH∗(M,ω): \emph{spectral length functional}, on the space of "generalized paths" in Ham(M,ω), closely related to both the Hofer length functional and spectral invariants and establish some of its properties. This functional is smooth on its…
We introduce a new family of matrix norms, the "local max" norms, generalizing existing methods such as the max norm, the trace norm (nuclear norm), and the weighted or smoothed weighted trace norms, which have been extensively used in the literature as regularizers for matrix reconstruction problems. We show that this…
Efficient algorithms for clustered Lasso and OSCAR reduce computational costs.
problem High dimensional regression with feature clustering.
method Efficient path algorithms for clustered Lasso and OSCAR, reducing computational costs.
result Proposed algorithms are more efficient than existing methods in numerical experiments.
We use the criteria of Lalonde and McDuff to determine a new class of examples of length minimizing paths in the group Ham(M). For a compact symplectic manifold M of dimension two or four, we show that a path in Ham(M), generated by an autonomous Hamiltonian and starting at the identity, which induces no non-cons…
This paper analyzes shallow ReLU networks in L^p and Sobolev spaces, focusing on approximation and generalization.
problem Approximation and generalization of shallow ReLU networks in L^p and Sobolev spaces.
method Spherical harmonic analysis and embeddings into spectral Barron spaces for L^p spaces, path-norm control for Sobolev spaces.
result Minimax-optimal rates for nonparametric regression with shallow ReLU networks under path-norm control.
Recently theoretical guarantees have been obtained for matrix completion in the non-uniform sampling regime. In particular, if the sampling distribution aligns with the underlying matrix's leverage scores, then with high probability nuclear norm minimization will exactly recover the low rank matrix. In this article, we…
The ℓ-1 norm based optimization is widely used in signal processing, especially in recent compressed sensing theory. This paper studies the solution path of the ℓ-1 norm penalized least-square problem, whose constrained form is known as Least Absolute Shrinkage and Selection Operator (LASSO). A solution path …
This paper concerns model reduction of dynamical systems using the nuclear norm of the Hankel matrix to make a trade-off between model fit and model complexity. This results in a convex optimization problem where this trade-off is determined by one crucial design parameter. The main contribution is a methodology to app…
Constructs weight 1/2 multiplier systems for a specific group and relates to geometric edge paths.
problem Constructing weight 1/2 multiplier systems for a specific group.
method Defines an eta function and Rademacher symbol, relates to geometric edge paths in a triangulation of the upper half plane.
result Relates weight 1/2 multiplier systems to geometric edge paths.
Sharp inequality on Siegel domain involving weighted norms and sub-Laplacian.
problem Establishing a Sobolev trace inequality on a specific domain.
method Using weighted norms and fractional powers of sub-Laplacian on Heisenberg group.
result Sharp Sobolev trace inequality on Siegel domain involving weighted norms.
New algorithms avoid weight transport, outperforming current deep learning methods.
problem Current deep learning algorithms rely on weight transport, which is biologically implausible.
method Two mechanisms: weight mirror and modified Kolen-Pollack algorithm, using random feedback weights.
result These mechanisms outperform feedback alignment and other methods on visual recognition tasks.
Lie groups with bi-invariant distance are products of abelian and compact groups.
problem Characterizing Lie groups with bi-invariant distances.
method Analyzing the structure of Lie groups and introducing a Finsler norm.
result The sectional curvature of bi-invariant distances is non-negative and vanishes only for abelian subalgebras.
Gradient flow on softmax attention minimizes nuclear norm of weight matrices.
problem Classification with separate key and query weight matrices.
method Gradient flow on exponential loss, separability assumption, reparameterization, approximate KKT conditions.
result Gradient flow implicitly minimizes nuclear norm of weight matrices, contrasting with Frobenius norm minimization.
This work shows how penalising bias terms in norm regularisation leads to sparse solutions.
problem Understanding the relation between parameter norm regularization and the sparsity of neural network solutions.
method Analyzes one hidden ReLU layer networks with unidimensional data, showing the norm required for function representation and the importance of the bias term's norm.
result Penalising the bias terms in regularisation leads to sparse solutions, enforcing the uniqueness and sparsity of the minimal norm interpolator.
Fine-tunes deep neural networks to match theoretical bounds on generalization errors.
problem Improve generalization errors of deep neural networks by constraining weight norms.
method Proposes a two-stage renormalization procedure and a fine-grained SGD algorithm for training DNNs with constrained weights.
result Empirical generalization errors of DNNs are closer to theoretical bounds, improving accuracy.
We provide rigorous guarantees on learning with the weighted trace-norm under arbitrary sampling distributions. We show that the standard weighted trace-norm might fail when the sampling distribution is not a product distribution (i.e. when row and column indexes are not selected independently), present a corrected var…
New framework explains deep neural networks using variational spline theory.
problem Understanding functions learned by deep neural networks.
method Developed a variational framework and function space.
result Deep ReLU networks are solutions to regularized data fitting problems over the proposed function space.
The paper proposes a method to improve random forest classification accuracy by weighting trees based on their decision path reliability.
problem Random forests' uniform voting fails to correct errors in regions where incorrect tree representations outnumber correct ones.
method The paper introduces using the structural pattern of each tree's decision path as an instance-adaptive reliability signal to identify and weight more reliable trees.
result Using the proposed method yields a statistically significant accuracy improvement over RF on 36 binary classification benchmarks.
New MCMC method improves sampling from multimodal distributions.
problem Sampling from multimodal distributions is challenging for classical MCMC methods.
method Interpolating along the diffusion path, preserving mode weights and mixing properties.
result MAD-Path sampler improves global exploration and mode-weight estimation.
This is the fourth article of our series. Here, we study weighted norm inequalities for the Riesz transform of the Laplace-Beltrami operator on Riemannian manifolds and of subelliptic sum of squares on Lie groups, under the doubling volume property and Gaussian upper bounds.
Weight normalization and reparametrized gradient descent adaptively regularize weights and converge to minimum l2 norm solutions.
problem Adapting to non-convex weight normalization for convergence to minimum l2 norm solutions.
method Weight normalization and reparametrized projected gradient descent (rPGD) for overparametrized least-squares regression.
result rPGD converges close to the minimum l2 norm solution, even for far-from-zero initializations.
This paper studies neural networks with bounded norms to avoid the curse of dimensionality.
problem The curse of dimensionality in approximating functions by neural networks.
method Investigates over-parameterized two-layer neural networks with norm constraints in RKHS.
result Improved sample complexity and generalization bounds for neural networks with bounded norms.
Characterizes inductive bias in multi-channel linear CNNs with bounded weight norm.
problem Understanding the inductive bias in multi-channel linear convolutional networks.
method Function space characterization and empirical testing of gradient descent.
result The inductive bias depends on the number of output channels for multi-channel inputs but not for single-channel inputs.
Any Sasakian structure can be closely mimicked by embeddings into weighted spheres.
problem Approximating Sasakian structures on closed manifolds.
method Using CR embeddings into weighted Sasakian spheres and strengthening previous approximation results.
result Sasakian structures can be approximated in the Cq-norm by embeddings into weighted Sasakian spheres. URGE improves diffusion model quality without gradients or Hessian.
problem Improving sample quality in diffusion models without gradient evaluations.
method Path-wise importance reweighting via Girsanov change of measure.
result URGE achieves better generation quality than existing methods.
Consider a weighted or unweighted k-nearest neighbor graph that has been built on n data points drawn randomly according to some density p on R^d. We study the convergence of the shortest path distance in such graphs as the sample size tends to infinity. We prove that for unweighted kNN graphs, this distance converges …
A new framework uses matrix flows to unify frequentist and Bayesian approaches for sparse GGMs.
problem Challenges in studying conditional independence among many variables with few observations.
method General framework for variational inference with matrix-variate Normalizing Flow in Gaussian Graphical Models.
result Unified benefits of frequentist and Bayesian frameworks for sparse GGMs.