Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

117234350467 · Jun 202019922001200920172026
48 results for large width

Lectures on deep learning properties in infinite and large-width networks.

problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.

This paper studies large-width asymptotics for ReLU neural networks with α-Stable initializations.

problem Characterizing the large-width behavior of ReLU neural networks with α-Stable initializations.
method Analysis of the large-width distributions and training dynamics of ReLU neural networks initialized with α-Stable distributions.
result For ReLU neural networks with α-Stable initializations, the large-width training dynamics achieve zero training error at a linear rate, characterized by a random kernel.

We extend the classical definition of {\it width} to higher dimensional, smooth codimension 2 knots and show in each dimension there are knots of arbitrarily large width.

2019-02-19abs ↗pdf ↗

Large learning rates work surprisingly well in standard parameterization, contrary to theory.

problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.

Wide neural networks can degrade performance, contrary to conventional wisdom.

problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.

The macroscopic version of Urysohn width for scalar curvature is disproven in high dimensions.

problem Disproving the macroscopic version of Gromov's Urysohn width conjecture for scalar curvature.
method Novel estimate on Urysohn width of circle bundles and a new notion of ruling for Riemannian manifolds.
result The macroscopic version of Gromov's Urysohn width conjecture for scalar curvature is false in dimensions four and above.

This work challenges the Neural Tangent Kernel's role in overparameterized neural networks, especially with large width and depth.

problem The Neural Tangent Kernel's behavior in overparameterized neural networks with large width and depth is unclear.
method Experimental and theoretical analysis of ReLU networks with large width and depth.
result The aggregate norm of hidden neuron deviations does not vanish in infinitely-wide ReLU networks, indicating non-trivial behavior.

Study on neural network initialization with shaped infinite depth-and-width networks.

problem Understanding the distribution of random covariance matrices in shaped infinite-depth-and-width networks.
method Introduced the Neural Covariance SDE to model the distribution of the random covariance matrix.
result Identified the precise scaling of the activation function necessary for a non-trivial limit.

New optimizers control network width scaling, improving stability and transfer across different model sizes.

problem Designing stable optimizers for networks of varying widths.
method Interpreting optimizers as steepest descent under mean-normalized operator norms, enabling layerwise composability and width-independent bounds.
result New optimizers like row normalization and column normalization provide stable learning-rate transfer across different model widths.

Study on 1-Uryson width of polyhedra and their covers.

problem Existence of Riemannian polyhedra with bounded 1-Uryson width of covers but unbounded in the polyhedron itself.
method Investigated specific cases of virtually cyclic fundamental groups and Riemannian surfaces, showing bounds on 1-Uryson width.
result For compact polyhedra with virtually cyclic fundamental groups, 1-Uryson width of polyhedron is bounded by that of its universal cover.

Quantitative CLTs show neural network distributions converge to Gaussian as width increases.

problem Understanding the distribution of fully connected neural networks with random weights and biases.
method Analyzing the distribution of a fully connected neural network with random Gaussian weights and biases, proving quantitative bounds on normal approximations.
result The distance between a random fully connected network and the corresponding infinite width Gaussian process scales like nγn^{-γ} for γ>0γ>0.

Study on fluctuations in neural network kernels and predictions, focusing on finite width effects.

problem Characterizing fluctuations in finite width neural networks.
method Dynamical mean field theory analysis of wide but finite feature learning neural networks.
result Fluctuations in kernels and predictions are dynamically coupled, leading to reduced variance in feature learning regimes.

We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation is exponential in the ratio of network depth to width. Thus, even in the limit of infinite overparameterization, the NTK is not determinist…

2019-09-13abs ↗pdf ↗

The study examines spectral dynamics in deep neural networks, predicting how outliers evolve during training.

problem Understanding spectral evolution in deep neural networks during training.
method Developed a two-level dynamical mean-field theory (DMFT) to track spectral dynamics.
result The theory predicts how outliers evolve with training time, width, output scale, and initialization variance.

Empirical study compares wide neural networks to kernel methods, resolving open questions.

problem Understanding the relationship between wide neural networks and kernel methods.
method Large-scale empirical study using various neural network architectures and kernel methods.
result Wide neural networks outperform fully-connected finite-width networks in some cases, but underperform convolutional finite-width networks.

In this work we construct a sequence of Riemannian metrics on the three-sphere with scalar curvature greater than or equal to 66 and arbitrarily large widths. Our procedure is based on the connected sum construction of positive scalar curvature metrics due to Gromov and Lawson. We develop analogies between the area of…

2015-03-08abs ↗pdf ↗

The paper proves conditions for the existence of small Urysohn width hypersurfaces in manifolds with positive scalar curvature.

problem Conditions for the existence of small Urysohn width hypersurfaces in manifolds with positive scalar curvature.
method Adaptation of Guth's macroscopic version of the Schoen-Yau descent argument.
result A complete Riemannian manifold with positive macroscopic scalar curvature contains a non-nullhomologous hypersurface of small Urysohn width.

Bayesian neural networks learn efficiently at infinite width, matching polynomial-width performance.

problem Understanding the inductive bias of infinite-width neural networks.
method Analyzing the reduced entropy and using subsampling techniques.
result The Bayesian mean-field learner generalizes exactly on polynomially-bounded targets.

Understanding the asymptotic behavior of wide networks is of considerable interest. In this work, we present a general method for analyzing this large width behavior. The method is an adaptation of Feynman diagrams, a standard tool for computing multivariate Gaussian integrals. We apply our method to study training dyn…

2019-09-25abs ↗pdf ↗

How large can be the width of Riemannian three-spheres of the same volume in the same conformal class? If a maximum value is attained, how does a maximising metric look like? What happens as the conformal class changes? In this paper, we investigate these and other related questions, focusing on the context of Simon-Sm…

2018-09-10abs ↗pdf ↗

Minimum width for ReLU networks on compact domain is exactly max{d_x, d_y, 2}

problem Characterizing the minimum width for ReLU networks to approximate functions on compact domains
method Analyzing the minimum width for LpL^p approximation of LpL^p functions from [0,1]d[0,1]^d to Rdy\mathbb R^{d_y} using ReLU-like activation functions
result The minimum width for LpL^p approximation on a compact domain is exactly max{d_x, d_y, 2} for ReLU-like activation functions

Given a 2-dimensional surface M and a constant C we construct a Riemannian metric g, so that diameter diam(M,g)=1 and every 1-cycle dividing M into two regions of equal area has length >C. It follows that there exists no universal inequality bounding 1-width of M in terms of its diameter. This answers a question of Ste…

2013-07-08abs ↗pdf ↗

Distributed implementations of mini-batch stochastic gradient descent (SGD) suffer from communication overheads, attributed to the high frequency of gradient updates inherent in small-batch training. Training with large batches can reduce these overheads; however, large batches can affect the convergence properties and…

2018-06-11abs ↗pdf ↗

NTK theory fails to predict practical behavior of large-width neural networks.

problem Theoretical limits of NTK do not match practical neural network architectures.
method Empirical investigation of NTK's applicability to large-width architectures.
result Practically relevant behavior of large-width architectures differs from NTK theory.

New approach predicts generalization of deep neural networks in proportional-width regime.

problem Predicting generalization of deep neural networks in proportional-width regime.
method Equivalent Wishart Ansatz for hierarchical empirical kernels, renormalized NNGP kernel.
result Renormalized NNGP kernel captures dominant stochastic fluctuations in deep neural networks.

The paper studies deep neural networks with Gaussian weights and finds their asymptotic behavior.

problem Understanding the behavior of deep neural networks with large width.
method Function-space perspective, Gaussian process analysis, weak convergence in large-width limit.
result Deep neural networks with large width converge to a continuous Gaussian process.

Study shows polynomial-width neural networks can closely approximate infinite-width networks in polynomial time.

problem Approximating dynamics of polynomial-width neural networks with infinite-width networks.
method Bounding approximation gap through a differential equation governed by mean-field dynamics, considering local Hessian.
result Polynomially many neurons are sufficient to closely approximate mean-field dynamics.

Gradient methods improve deep network training with tighter bounds and faster convergence.

problem Improving convergence and generalization of gradient methods for neural networks.
method Algorithmic stability analysis and novel bounds on excess risk.
result Gradient descent achieves optimal excess risk for deep nets with polynomial width conditions.

Study of deep linear neural networks with proportional width and depth.

problem Lack of descriptive power in Gaussian limit of deep linear neural networks.
method Proportional infinite-width infinite-depth limit for deep linear neural networks.
result Characterization of limiting distribution as a nontrivial mixture of Gaussians.

Study shows how large neural networks avoid overfitting through decoupling of feature learning and complexity growth.

problem Understanding inductive bias and generalization in large neural networks.
method Dynamical mean field theory applied to large two-layer networks.
result Training dynamics of large networks exhibit a separation of timescales, decoupling feature learning and overfitting.

Wide networks with polynomial activations have proven asymptotic behavior.

problem Understanding the behavior of neural networks in the large width limit.
method Proving a conjecture for deep networks with polynomial activation functions.
result Tight bounds on the behavior of wide networks during stochastic gradient descent and derivation of their finite-width dynamics.

New analysis shows optimal embedding learning rate depends on vocabulary size, not just model width.

problem Optimal learning rate for language model embeddings is not well understood, especially with large vocabularies.
method Theoretical analysis of training dynamics, interpolation between μμP and LV regimes.
result Optimal embedding learning rate scales as Θ(width)Θ(\sqrt{width}) in the LV regime, not Θ(width)Θ(width) as μμP predicts.

Study on spherical bodies of constant width on the unit sphere, proving bounds on their relative effective radius.

problem Understanding the smallest spherical bodies of constant width on the unit sphere.
method Analyzing spherical bodies of constant width on the unit sphere, constructing examples and applying geometric arguments.
result Proved non-trivial bounds on the relative effective radius of spherical bodies of constant width.

Maximal initial learning rate for deep ReLU networks identified.

problem Finding the optimal initial learning rate for deep neural networks.
method Simple approach to estimate maximal initial learning rate ηη^{\ast}, analyzing its behavior in constant-width fully-connected ReLU networks.
result Maximal initial learning rate ηη^{\ast} is well predicted as a power of depth × width, with specific conditions for network width and input layer training.

New Transformer architecture prevents rank degeneracy in deep attention models.

problem Rank degeneracy in deep attention models.
method Modified Softmax-based attention model with skip connections, centered at identity, and scaled logits.
result Existence of a stable SDE implies well-behaved covariance structure, preventing rank degeneracy.

The study analyzes deep linear networks from random initialization, capturing dynamics and hyperparameter effects.

problem Understanding training dynamics in deep linear networks from random initialization.
method Theoretical analysis of gradient descent dynamics in deep linear networks with random initialization and large data.
result Captures the 'wider is better' effect and hyperparameter transfer effects, contrasting with neural-tangent parameterization.

In order to investigate the origin of large price fluctuations, we analyze stock price changes of ten frequently traded NASDAQ stocks in the year 2002. Though the influence of the trading frequency on the aggregate return in a certain time interval is important, it cannot alone explain the heavy tailed distribution of …

2006-06-18abs ↗pdf ↗

Deep networks can perfectly classify two low-dimensional manifolds on a sphere with large depth and width.

problem Binary classification of two low-dimensional submanifolds on a sphere.
method Analysis of a deep fully-connected neural network trained to separate two submanifolds of the unit sphere.
result Randomly-initialized gradient descent can perfectly classify the two manifolds with high probability when the network depth is large relative to certain geometric and statistical properties of the data.

Analyzes DNNs trained with noisy gradients, finding FWCs negligible for large n.

problem Analyzing DNNs trained with noisy gradients.
method Introduced analytical framework to analyze non-Gaussian stochastic process.
result FWCs negligible for large n, improving CNN performance.

This work studies fluctuation in multilayer neural networks using mean field theory.

problem Understanding fluctuation in multilayer neural networks with mean field training.
method Developed a second-order mean field limit to capture fluctuation, demonstrating stability of gradient descent training.
result Gradient descent training in multilayer networks biases towards minimal fluctuation, even after convergence.

Residual networks with depthwise hyperparameter scaling transfer optimal hyperparameters across width and depth.

problem The challenge of hyperparameter tuning in deep learning, especially for large models.
method Combining μμP parameterization with residual networks having a residual branch scale of 1/extdepth1/\sqrt{ ext{depth}}.
result Optimal hyperparameters transfer across width and depth in residual networks trained with this parameterization.

This paper analyzes deep Stable neural networks, showing convergence rates under different growth settings.

problem Analyzing the behavior of deep Stable neural networks as width increases.
method Large-width asymptotic analysis and convergence rates for fully connected feed-forward deep Stable NNs.
result The rescaled deep Stable NN converges weakly to a Stable SP under joint growth, with sup-norm convergence rates established.