Study shows upper limit for torical band width with spectral curvature bounds.
problem Understanding the band width of torical bands with spectral curvature constraints.
method Used the warped \( μ\)-bubble method with spectral curvature bounds.
result Upper bound for the band width of torical bands is established.
Derives generalizations of the long neck principle and spectral width inequality.
problem Understanding the spectral width of geodesic collar neighborhoods.
method Spinorial Callias operator approach and relative Gromov-Lawson pair.
result Generalizations of the long neck principle and spectral width inequality.
Unified spectral framework for μP under joint width-depth scaling.
problem Challenges in stable feature learning and HP transfer for width-depth scaled models.
method Developed a simple and unified spectral framework for μP under joint width-depth scaling.
result Unified and generalized μP formulation for practical architectures with multi-transformation branches.
The study examines spectral dynamics in deep neural networks, predicting how outliers evolve during training.
problem Understanding spectral evolution in deep neural networks during training.
method Developed a two-level dynamical mean-field theory (DMFT) to track spectral dynamics.
result The theory predicts how outliers evolve with training time, width, output scale, and initialization variance.
The paper explores rigidity theorems for spectral curvature bounds in 3-manifolds.
problem Classical rigidity results in scalar curvature geometry are extended to the spectral setting.
method Warped μ μ μ -bubble method is systematically employed to classify stable weighted minimal hypersurfaces and establish band width estimates. result Classification theorems and band width estimates for spectral Ricci and scalar curvatures are proven.
Extends spectral torus band inequalities for compact manifolds with scalar curvature bounds.
problem Proving upper bounds for the width of compact manifolds with boundary.
method Utilizes spacetime harmonic functions, μ-bubbles, and spinorial Callias operators.
result Generalizes Schoen-Yau black hole existence theorem to higher dimensions.
Geodesic balls with non-negative Ricci curvature have a sharp lower bound on their first Dirichlet eigenvalue.
problem Finding a sharp lower bound for the first Dirichlet eigenvalue of geodesic balls.
method Quantitative explicit inequality linking the width of geodesic balls to the spectral gap.
result A quantitative inequality relating the width of geodesic balls to the spectral gap between the first Dirichlet eigenvalue and its lower bound.
Investigates spectral properties of neural networks, showing invariance under certain conditions.
problem Understanding the spectral evolution and invariance in linear-width neural networks.
method Empirical and theoretical analysis of spectra of weight matrices in high-dimensional settings.
result Spectra of weight matrices are invariant under certain training conditions, with implications for feature learning.
Lobb observed in [arXiv:1103.1412] that each equivariant sl(N) Khovanov-Rozansky homology over C[a] admits a standard decomposition of a simple form. In the present paper, we derive a formula for the corresponding Lee-Gornik spectral sequence in terms of this decomposition. Based on this formula, we give a simple alter…
Develops a new theory for neural systems stability and width effects.
problem Stability and finite-width effects in deep neural systems.
method Gauge-covariant stochastic effective field theory using classical commuting fields.
result Predicts the edge of chaos and low-frequency spectral deformation.
Define mass invariant for asymptotically hyperbolic 3-manifolds with toroidal infinity.
problem Define mass invariant for asymptotically hyperbolic 3-manifolds with toroidal infinity.
method Define mass invariant for asymptotically hyperbolic 3-manifolds with toroidal infinity.
result Define mass invariant for asymptotically hyperbolic 3-manifolds with toroidal infinity.
Deep networks can be biased to learn top eigenfunctions of the kernel outside the training set.
problem Spectral bias of deep networks in the kernel regime.
method Quantitative bounds on L 2 L^2 L 2 difference between finite-width and infinite-width network trajectories. result Deep networks learn top eigenfunctions of the Neural Tangent Kernel over the entire input space, not just the training set.
The paper proves rigidity for certain product spaces and bounds for band widths.
problem Proving rigidity for product spaces and bounds for band widths.
method Combining stable weighted slicing with a spectral Dirac operator argument.
result Closed spin ( M n , g ) (M^n,g) ( M n , g ) is isometrically covered by S n − m i m e s R m S^{n-m} imes\mathbb{R}^m S n − m im es R m under certain conditions. Paper connects neural networks to Gaussian processes for understanding double-descent.
problem Understanding the double-descent phenomenon in neural networks.
method Uses techniques from random matrix theory and Gaussian processes.
result Establishes a connection between NNGP and random matrix theory for neural networks.
Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
Study how generalization scales with model size and data in quadratic neural networks.
problem Understanding how generalization scales with model size and data in quadratic neural networks.
method Analyzed ℓ 2 \ell_2 ℓ 2 -regularized empirical test error minimization in a quadratic two-layer network with finite-sample setting and structured data. result Revealed a phase diagram with distinct scaling regimes as the number of parameters varies, showing data-dependent power laws controlled by spectral structure of the target.
Spectral analysis shows neural networks separate from linear methods in approximating functions.
problem Separating two-layer neural networks from linear methods in function approximation.
method Spectral-based approach using Kolmogorov width and kernel spectrum.
result Upper and lower bounds on separation, explicit hard functions identified.
New algorithms reduce dueling bandits' regret with neural networks and efficient exploration.
problem Optimizing dueling bandits with neural networks for better performance.
method Combines shallow exploration strategies with neural networks for utility approximation, using iterative self-improvement and spectral analysis to reduce network width.
result Achieves sublinear regret of O ~ ( d ∑ t = 1 T σ t 2 + d T ) \widetilde{\mathcal{O}}(d\sqrt{\sum_{t=1}^{T} σ_t^2} + \sqrt{dT}) O ( d ∑ t = 1 T σ t 2 + d T ) . Given a complete non-compact surface embedded in R^3, we consider the Dirichlet Laplacian in a layer of constant width about the surface. Using an intrinsic approach to the layer geometry, we generalise the spectral results of an original paper by Duclos et al. to the situation when the surface does not possess poles. …
Proposes a deep network for multi-class classification using spectral training and Gaussian kernel.
problem Multi-class classification with deep networks.
method Spectral training with linear weights and Gaussian kernel activation, constrained on Stiefel Manifold.
result Theoretical guarantee of global optimum and insight into network generalization.
Study on neural network dynamics in high dimensions with quadratic activation.
problem Understanding training dynamics in overparameterized neural networks.
method Derivation of gradient flow equations and analysis under l2-regularization.
result Characterization of estimator performance and spectral properties in the high-dimensional limit.
Study shows deterministic equivalent for neural network kernel convergence.
problem Understanding convergence of neural network kernels.
method Analyzes empirical spectral distribution of Conjugate Kernel, proving convergence to a deterministic limit.
result Obtains a deterministic equivalent for the Stieltjes transform and resolvent of the Conjugate Kernel.
TopoNTK kernel captures higher-order interactions in simplicial complexes.
problem Graph neural networks miss higher-order interactions in relational systems.
method Introduces TopoNTK, an infinite-width kernel for simplicial message passing.
result TopoNTK captures topology invisible to graph kernels, improving expressivity and interpretability.
This study analyzes why attention layers in neural networks can cause signal loss and proposes a solution.
problem Pathological behavior of attention layers in neural networks, leading to signal loss.
method Spectral analysis using Random Matrix Theory to identify and mitigate rank collapse in width.
result A novel solution to mitigate rank collapse in width by removing outlier eigenvalues.
Gradient descent converges linearly in finite-width networks with positive NTK and compatible conditions.
problem Local convergence of gradient descent in finite-width networks.
method Positive Neural Tangent Kernel (NTK), local Polyak-Łojasiewicz inequality, fixed-step containment in Locally Quasi-Convex Region (LQCR).
result Linear convergence achieved under specific conditions.
By adopting Multifractal detrended fluctuation (MF-DFA) analysis methods, the multifractal nature is revealed in the high-frequency data of two typical indexes, the Shanghai Stock Exchange Composite 180 Index (SH180) and the Shenzhen Stock Exchange Composite Index (SZCI). The characteristics of the corresponding multif…
Linearized attention fails to converge to NTK limit even at large widths.
problem Understanding the convergence of attention mechanisms to the kernel regime.
method Analyzes linearized attention and its relationship to the NTK limit, considering practical widths and conditions.
result Linearized attention does not converge to its NTK limit at any practical width, revealing a fundamental trade-off.
Analyzes how diffusion models learn, revealing a spectral bias in structure mastery.
problem Understanding the learning dynamics and bias in diffusion models.
method Developed an analytical framework using a Gaussian-equivalence principle to solve gradient-flow dynamics and integrate probability-flow ODEs.
result Exposes a universal inverse-variance spectral law: high-variance structure is mastered faster than low-variance detail.
Study shows neural operators can efficiently solve complex reaction-diffusion systems.
problem Efficiently solving nonlinear reaction-diffusion systems using neural operators.
method Laplacian-based neural operators applied to a generalized Gierer-Meinhardt system.
result Explicit approximation error bounds established for neural operators in terms of network parameters.
The paper explains generalization in kernel regression and deep neural networks using spectral bias and task-model alignment.
problem Understanding generalization in machine learning models, especially deep neural networks.
method Analytical expression for generalization error derived from statistical mechanics, applied to various kernels and data distributions.
result Spectral bias and task-model alignment explain generalization in kernel regression and deep neural networks.
The generalization error of deep neural networks via their classification margin is studied in this work. Our approach is based on the Jacobian matrix of a deep neural network and can be applied to networks with arbitrary non-linearities and pooling layers, and to networks with different architectures such as feed forw…
An algorithm calculates Gabai width for thousands of knots.
problem Calculating Gabai width for many knots.
method Algorithmic definition of Wirtinger width leading to efficient Gabai width bounds.
result Proved Wirtinger width equals Gabai width for knots.
Recently, a number of works have studied clustering strategies that combine classical clustering algorithms and deep learning methods. These approaches follow either a sequential way, where a deep representation is learned using a deep autoencoder before obtaining clusters with k-means, or a simultaneous way, where dee…
New insights into neural network feature learning through multi-step gradient descent.
problem Understanding feature learning in two-layer neural networks with limited width.
method Characterization of feature learning through two steps of gradient descent with specific step sizes.
result The second step of gradient descent reveals multiple learned directions, not limited to a single direction as in the first step.
Study on neural networks' sample complexity with one hidden layer.
problem Understanding how sample complexity is affected by network architecture and norm constraints.
method Norm-based uniform convergence bounds for scalar-valued one-hidden-layer networks, focusing on spectral and Frobenius norms.
result Spectral norm control is insufficient for uniform convergence guarantees, but Frobenius norm control is sufficient, with conditions.
The angular power spectrum characterizes neural network complexity.
problem Characterizing the complexity of deep neural networks.
method Using the angular power spectrum of the limiting field to characterize network complexity.
result Classified neural networks as low-disorder, sparse, or high-disorder.
Two new algorithms improve robust PCA and Schatten packing.
problem Robustly estimating the top eigenvector of corrupted sub-Gaussian data.
method Two iterative filtering and nearly-linear time algorithms.
result First polynomial-time algorithms for non-trivial covariance estimation.
Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.
problem Comparing BNNs with different widths due to conflicting model properties and inference intractability.
method Empirical comparison of finite- and infinite-width BNNs, analyzing performance under model mismatch.
result Increasing width can hurt BNN performance when the model is mis-specified, and finite-width BNNs generalize better under model mismatch.
The isospectral problem for p-widths is solved using Zoll metrics on S^2.
problem Determine if a Riemannian manifold is uniquely determined by its p-widths.
method Construct counterexamples on S^2 using Zoll metrics and properties of geodesic p-widths.
result Many counterexamples exist on S^2, showing uniqueness is not guaranteed.
Width trees link link invariants and bridge number.
problem Understanding link invariants through geometric structures.
method Associate width trees to links and use their geometric properties to bound link invariants.
result Width trees uniquely realize certain link invariants under specific conditions.
The paper proves geometric and spectral alignment for deep neural networks.
problem Understanding the singular spectra of deep neural network layers.
method Proves deterministic quotient-geometric estimates for singular spectra of Frobenius-normalized layer factors.
result Exact power-law spectra form a trace-normalized Cartan orbit under Frobenius normalization.
Lectures on deep learning properties in infinite and large-width networks.
problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.
Computed p-widths for hemisphere, first for manifolds with boundary.
problem Finding p-widths for manifolds with boundary.
method Computed p-widths for the hemisphere.
result First known p-widths for a manifold with boundary.
Polygon p p p -widths are found via billiard trajectories.
problem Finding p p p -widths of polygons. method Proved via billiard trajectories and computed specific cases.
result Polygon p p p -widths are achieved by billiard trajectories. A number of results for C 2 ^2 2 -smooth surfaces of constant width in Euclidean 3-space E 3 {\mathbb{E}}^3 E 3 are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…
Computed p-widths for real projective plane.
problem Calculating p-widths for real projective plane.
method Standard metric used to compute p-widths.
result Computed p-widths for real projective plane.
Study bounds Urysohn width of manifolds under surgeries.
problem Bounding Urysohn width of manifolds after surgeries.
method Analyzes connected sums and universal covers, applies to general surgeries.
result Optimal constants in estimates of width bounds are shown.
Residual networks with block width max(d_x, d_y) approximate all functions.
problem Achieving universal approximation with residual networks.
method Established bounds on block width for different activation functions.
result Minimum block width for universal approximation is max(d_x, d_y) with inner width 1.