Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

8.3%16.7%25.0%33.3% · Jul 199219922001200920172026
48 results for width definition

Wide CNNs outperform infinite width networks, revealing scaling laws.

problem Understanding the performance difference between finite and infinite width convolutional networks.
method Diagrammatic approach to derive asymptotic width dependence for various quantities.
result The difference in performance between finite and infinite width models vanishes at a definite rate with respect to model width.

Minimum width for ReLU networks to approximate L^p functions is max(d_x+1, d_y).

problem Characterizing the minimum width for ReLU networks to approximate L^p functions.
method Analyzing networks with ReLU activation functions and proving the minimum width required.
result The minimum width required for the universal approximation of L^p functions is exactly max(d_x+1, d_y).

Motivated by M. Scharlemann and A. Thompson's definition of thin position of 3-manifolds, we define the width of a handle decomposition a 4-manifold and introduce the notion of thin position of a compact smooth 4-manifold. We determine all manifolds having width equal to {1,,1}\{1,\dots, 1\}, and give a relation between th…

2018-05-22abs ↗pdf ↗

This paper addresses the so-called conformal capacities in Rn\mathbb R^n, n3n\ge 3, through comparing three existing definitions (due to Betsakos, Colesanti-Cuoghi, Anderson-Vamananmurthy-Fuglede respectively) and studying their associated iso-capacitary inequalities with connection to half-diameter, mean-width, mean-c…

2013-09-14abs ↗pdf ↗

The present paper proposes generalized Gaussian kernel adaptive filtering, where the kernel parameters are adaptive and data-driven. The Gaussian kernel is parametrized by a center vector and a symmetric positive definite (SPD) precision matrix, which is regarded as a generalization of the scalar width parameter. These…

2018-04-25abs ↗pdf ↗

Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.

problem Comparing BNNs with different widths due to conflicting model properties and inference intractability.
method Empirical comparison of finite- and infinite-width BNNs, analyzing performance under model mismatch.
result Increasing width can hurt BNN performance when the model is mis-specified, and finite-width BNNs generalize better under model mismatch.

Lobb observed in [arXiv:1103.1412] that each equivariant sl(N) Khovanov-Rozansky homology over C[a] admits a standard decomposition of a simple form. In the present paper, we derive a formula for the corresponding Lee-Gornik spectral sequence in terms of this decomposition. Based on this formula, we give a simple alter…

2012-11-28abs ↗pdf ↗

New model classes for function approximation by neural networks defined on domains.

problem Defining novel model classes for function approximation on bounded domains.
method Introducing weighted variation spaces to define new model classes on domains.
result New model classes are strictly larger than classical ones but maintain the same NNA rates.

Lectures on deep learning properties in infinite and large-width networks.

problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.

A number of results for C2^2-smooth surfaces of constant width in Euclidean 3-space E3{\mathbb{E}}^3 are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…

2007-04-24abs ↗pdf ↗

Residual networks with block width max(d_x, d_y) approximate all functions.

problem Achieving universal approximation with residual networks.
method Established bounds on block width for different activation functions.
result Minimum block width for universal approximation is max(d_x, d_y) with inner width 1.

Bayesian neural networks with dependent weights converge to Gaussian mixtures.

problem Limitations of standard Gaussian priors in neural networks.
method Posterior analysis with Gaussian likelihood for networks with dependent weights.
result Posterior distribution identified in the wide-width limit, ensuring invertibility of random covariance matrix.

While studying the existence of closed geodesics and minimal hypersurfaces in compact manifolds, the concept of width was introduced in different contexts. Generally, the width is realized by the energy of the closed geodesics or the volume of minimal hypersurfaces, which are found by the Minimax argument. Recently, Ma…

2016-12-20abs ↗pdf ↗

Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.

problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.

Study infinite-depth limits of neural networks with fixed width.

problem Understanding the behavior of neural networks as depth increases with fixed width.
method Analyzing finite-width residual networks with random Gaussian weights, focusing on the infinite-depth limit.
result The pre-activations converge to a zero-drift diffusion process, differing from the infinite-width limit.

In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…

2010-08-30abs ↗pdf ↗

Let MM be a Riemannian manifold with a smooth boundary. The main question we address in this article is: "When is the Laplace-Beltrami operator Δ ⁣:Hk+1(M)H01(M)Hk1(M)Δ\colon H^{k+1}(M)\cap H^1_0(M) \to H^{k-1}(M), kN0k\in \mathbb{N}_0, invertible?" We consider also the case of mixed boundary conditions. The study of this main question lea…

2016-11-01abs ↗pdf ↗

Wide neural networks can degrade performance, contrary to conventional wisdom.

problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.

A new complexity measure for neural networks improves upon classical methods.

problem Lack of a refined complexity measure for comparing different neural network architectures, especially permutation-invariant ones.
method Introduced an equivalence relation among linear functions and counted them relative to this relation.
result The new complexity measure clearly distinguishes between different models and increases exponentially with depth.

The study analyzes a three-layer neural network's training dynamics using a functional-space mean-field theory.

problem Understanding the training dynamics of partially-trained three-layer neural networks.
method Generalized mean-field theory to functional spaces, proving convergence and feature learning.
result The training loss of the model decays to zero at a linear rate in the L2L_2 regression setting.

Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.

problem Understanding the optimal balance between depth and width in self-attention models.
method Theoretical predictions and empirical ablations on networks of varying depths and widths.
result An optimal width of 30K is recommended for a 1-Trillion parameter network, marking a significant width for self-attention models.

We present an alternative proof of the following fact: the hyperspace of compact closed subsets of constant width in Rn\mathbb R^n is a contractible Hilbert cube manifold. The proof also works for certain subspaces of compact convex sets of constant width as well as for the pairs of compact convex sets of constant rela…

2004-01-07abs ↗pdf ↗

New framework for understanding infinite-width neural networks.

problem Understanding the infinite-width limit behavior of neural networks.
method General framework to study limit behavior of neural models based on hyperparameter scaling.
result Derives scaling for existing mean-field and neural tangent kernel limits and introduces new dynamically stable limits.

Study examines dependence properties of Bayesian neural network units in finite-width networks.

problem Understanding dependence properties of hidden units in practical finite-width Bayesian neural networks.
method Theoretical analysis and empirical evaluation of depth and width impacts.
result Hidden units in finite-width Bayesian neural networks are dependent, contrary to the infinite-width limit assumption.

We study surface knots in 4-space by using generic planar projections. These projections have fold points and cusps as their singularities and the image of the singular point set divides the plane into several regions. The width (or the total width) of a surface knot is a numerical invariant related to the number of po…

2009-05-21abs ↗pdf ↗

Improved standard parameterization yields well-defined neural tangent kernel.

problem Extrapolation of standard parameterization to infinite width is problematic.
method Proposed an improved extrapolation of the standard parameterization.
result Improved standard parameterization yields similar accuracy to NTK parameterization but with better correspondence to finite width networks.

Bayesian neural networks learn efficiently at infinite width, matching polynomial-width performance.

problem Understanding the inductive bias of infinite-width neural networks.
method Analyzing the reduced entropy and using subsampling techniques.
result The Bayesian mean-field learner generalizes exactly on polynomially-bounded targets.