This paper improves neural network approximation for analytic functions with adjustable depth and width.
problem Approximating analytic functions using neural networks with depth and width parameters.
method Characterizes approximation rates as a joint function of width (N) and depth (L) for ReLU networks.
result Establishes upper bounds for analytic function approximation rates of O(N^(-CL^τ)) with τ influenced by N and L.
Wider networks improve natural accuracy but worsen perturbation stability, affecting overall robustness.
problem Understanding the tradeoff between natural accuracy and perturbation stability in wider neural networks for adversarial robustness.
method Careful examination of the relationship between network width, robust regularization parameter λ, and perturbation stability using neural tangent kernels.
result Wider networks can achieve better natural accuracy but worse perturbation stability, leading to potentially worse overall model robustness.
Develops a robust hedging valuation adjustment measure for dynamic hedging under liquidity-demand stress.
problem Dynamic hedging under liquidity-demand stress
method Define robust HVA as the worst-case expected loss over a relative-entropy neighborhood of the loss distribution generated by simulated rebalancing and maturity-unwind trades.
result Distinguishes fixed-radius convention from fixed benchmark-stress convention and shows wider no-trade bands lower rebalancing costs but raise hedge-error risk.
Paper develops a robust HVA measure for dynamic hedging under liquidity stress.
problem Valuation of dynamic hedging under liquidity stress.
method Defines robust HVA as worst-case expected loss over a relative-entropy neighborhood of loss distributions for no-trade bands.
result Wider no-trade bands lower rebalancing costs but increase hedge-error risk.
A remarkable characteristic of overparameterized deep neural networks (DNNs) is that their accuracy does not degrade when the network's width is increased. Recent evidence suggests that developing compressible representations is key for adjusting the complexity of large networks to the learning task at hand. However, t…
Deep networks can approximate various activation functions with modest adjustments.
problem Expressive power of deep neural networks with diverse activation functions.
method Approximation of any activation function in set A by ReLU networks with specific scaling factors.
result Approximation of any activation function in a specific subset of A by ReLU networks with (1,1) scaling factors.
We define the Wirtinger width of a knot. Then we prove the Wirtinger width of a knot equals its Gabai width. The algorithmic nature of the Wirtinger width leads to an efficient technique for establishing upper bounds on Gabai width. As an application, we use this technique to calculate the Gabai width of approximately …
Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.
problem Comparing BNNs with different widths due to conflicting model properties and inference intractability.
method Empirical comparison of finite- and infinite-width BNNs, analyzing performance under model mismatch.
result Increasing width can hurt BNN performance when the model is mis-specified, and finite-width BNNs generalize better under model mismatch.
CNN predicts stock fluctuations using company news headlines.
problem Predicting next-day stock fluctuations based on company-specific news.
method Convolutional Neural Network (CNN) with reduced filter dimensions and multiple hidden layers. Fine-tuned word embeddings and various filter widths.
result 61.7% classification accuracy achieved using pre-learned embeddings.
The isospectral problem for p-widths is solved using Zoll metrics on S^2.
problem Determine if a Riemannian manifold is uniquely determined by its p-widths.
method Construct counterexamples on S^2 using Zoll metrics and properties of geodesic p-widths.
result Many counterexamples exist on S^2, showing uniqueness is not guaranteed.
Width trees link link invariants and bridge number.
problem Understanding link invariants through geometric structures.
method Associate width trees to links and use their geometric properties to bound link invariants.
result Width trees uniquely realize certain link invariants under specific conditions.
Lectures on deep learning properties in infinite and large-width networks.
problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.
Computed p-widths for hemisphere, first for manifolds with boundary.
problem Finding p-widths for manifolds with boundary.
method Computed p-widths for the hemisphere.
result First known p-widths for a manifold with boundary.
Polygon p-widths are found via billiard trajectories.
problem Finding p-widths of polygons. method Proved via billiard trajectories and computed specific cases.
result Polygon p-widths are achieved by billiard trajectories. A number of results for C2-smooth surfaces of constant width in Euclidean 3-space E3 are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…
Computed p-widths for real projective plane.
problem Calculating p-widths for real projective plane.
method Standard metric used to compute p-widths.
result Computed p-widths for real projective plane.
Study bounds Urysohn width of manifolds under surgeries.
problem Bounding Urysohn width of manifolds after surgeries.
method Analyzes connected sums and universal covers, applies to general surgeries.
result Optimal constants in estimates of width bounds are shown.
Residual networks with block width max(d_x, d_y) approximate all functions.
problem Achieving universal approximation with residual networks.
method Established bounds on block width for different activation functions.
result Minimum block width for universal approximation is max(d_x, d_y) with inner width 1.
While studying the existence of closed geodesics and minimal hypersurfaces in compact manifolds, the concept of width was introduced in different contexts. Generally, the width is realized by the energy of the closed geodesics or the volume of minimal hypersurfaces, which are found by the Minimax argument. Recently, Ma…
RAGIC predicts stock intervals with risk considerations, improving prediction accuracy and coverage.
problem Limited success in predicting stock market outcomes due to stochastic nature and risk oversight.
method RAGIC uses a GAN with a risk module and temporal module to generate risk-sensitive stock intervals.
result RAGIC achieves a consistent 95% coverage with narrow interval widths, balancing accuracy and risk.
Proves conjecture about sphere widths under rotational symmetry.
problem Width stability of rotationally symmetric metrics.
method Proof of conjecture and extensions to higher dimensions.
result Stability of min-max width under rotational symmetry.
New link invariants from diagram colorings match link widths.
problem Defining link widths via diagram colorings.
method Colorings of link diagrams to define invariants and prove their equivalence to link widths.
result Invariants of link widths calculated algorithmically.
Study infinite-depth limits of neural networks with fixed width.
problem Understanding the behavior of neural networks as depth increases with fixed width.
method Analyzing finite-width residual networks with random Gaussian weights, focusing on the infinite-depth limit.
result The pre-activations converge to a zero-drift diffusion process, differing from the infinite-width limit.
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
AdaLoss optimizes adaptive learning rates for efficient convergence in various models.
problem Efficiently optimizing adaptive learning rates for gradient descent methods.
method AdaLoss uses loss function information to dynamically adjust step sizes.
result AdaLoss achieves linear convergence in linear regression and robust global convergence in neural networks.
Sharp lower bound for first Neumann eigenvalue found in terms of diameter and width.
problem Finding the minimum value of the first Neumann eigenvalue for convex domains.
method Proved the sharp lower bound using diameter and width.
result Sharp lower bound for the first Neumann eigenvalue established.
We prove that among all constant width bodies of revolution, the minimum of the ratio of the volume to the cubed width is attained by the constant width body obtained by rotation of the Reuleaux triangle about an axis of symmetry.
Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
Wide neural networks can degrade performance, contrary to conventional wisdom.
problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.
Develops a new theory of width for embedded circles in Riemannian manifolds.
problem Defining and understanding the width of embedded circles in Riemannian manifolds.
method Morse-Lusternik-Schnirelmann theory applied to geodesics and minimising configurations.
result Classifies configurations of minimising geodesics intersecting embedded circles.
Paper proves a noncompact version of Gromov's band-width estimate.
problem Proving a precise upper bound for noncompact Riemannian bands.
method Developed a quantitative partitioned manifold index theory.
result Proved a version of Gromov's band-width estimate for noncompact Riemannian bands.
We discuss a possible definition for "k-width" of both a closed d-manifold Md, and on embedding Md↪eRn, n>d≥k, generalizing the classical notion of width of a knot. We show that for every 3-manifold 2-width(M3)≤2 but that there are embeddings $e_i: T^3 \hoo…
We extend the classical definition of {\it width} to higher dimensional, smooth codimension 2 knots and show in each dimension there are knots of arbitrarily large width.
Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.
problem Understanding the optimal balance between depth and width in self-attention models.
method Theoretical predictions and empirical ablations on networks of varying depths and widths.
result An optimal width of 30K is recommended for a 1-Trillion parameter network, marking a significant width for self-attention models.
We present an alternative proof of the following fact: the hyperspace of compact closed subsets of constant width in Rn is a contractible Hilbert cube manifold. The proof also works for certain subspaces of compact convex sets of constant width as well as for the pairs of compact convex sets of constant rela…
New framework for understanding infinite-width neural networks.
problem Understanding the infinite-width limit behavior of neural networks.
method General framework to study limit behavior of neural models based on hyperparameter scaling.
result Derives scaling for existing mean-field and neural tangent kernel limits and introduces new dynamically stable limits.
Wide CNNs outperform infinite width networks, revealing scaling laws.
problem Understanding the performance difference between finite and infinite width convolutional networks.
method Diagrammatic approach to derive asymptotic width dependence for various quantities.
result The difference in performance between finite and infinite width models vanishes at a definite rate with respect to model width.
Study examines dependence properties of Bayesian neural network units in finite-width networks.
problem Understanding dependence properties of hidden units in practical finite-width Bayesian neural networks.
method Theoretical analysis and empirical evaluation of depth and width impacts.
result Hidden units in finite-width Bayesian neural networks are dependent, contrary to the infinite-width limit assumption.
We study surface knots in 4-space by using generic planar projections. These projections have fold points and cusps as their singularities and the image of the singular point set divides the plane into several regions. The width (or the total width) of a surface knot is a numerical invariant related to the number of po…
Bayesian neural networks learn efficiently at infinite width, matching polynomial-width performance.
problem Understanding the inductive bias of infinite-width neural networks.
method Analyzing the reduced entropy and using subsampling techniques.
result The Bayesian mean-field learner generalizes exactly on polynomially-bounded targets.
New optimizers control network width scaling, improving stability and transfer across different model sizes.
problem Designing stable optimizers for networks of varying widths.
method Interpreting optimizers as steepest descent under mean-normalized operator norms, enabling layerwise composability and width-independent bounds.
result New optimizers like row normalization and column normalization provide stable learning-rate transfer across different model widths.
Smooth manifolds can be triangulated with graphs of bounded twin-width.
problem Understanding the structure of triangulations of smooth manifolds.
method Using Whitney's triangulation method and bounding the twin-width of specific graphs.
result Compact smooth manifolds have triangulations with graphs of bounded twin-width.
Study on ball widths and minimal submanifolds in space forms.
problem Understanding widths of balls and minimal submanifolds.
method Analyzing the area of equatorial balls and related bounds for minimal submanifolds.
result Lower bounds for the area of free boundary minimal submanifolds.
The paper proves rigidity theorems for area widths of Riemannian manifolds.
problem Characterizing metrics on Riemannian manifolds using their area widths.
method Analyzing the volume spectrum and spherical area widths.
result Rigidity theorems for specific metrics on projective spaces.
There are currently two parameterizations used to derive fixed kernels corresponding to infinite width neural networks, the NTK (Neural Tangent Kernel) parameterization and the naive standard parameterization. However, the extrapolation of both of these parameterizations to infinite width is problematic. The standard p…
The paper bounds the min-max width of embedded circles on spheres and manifolds.
problem Bounding the min-max width of embedded circles on spheres and manifolds.
method Inducing a sweepout by pairs of points in embedded circles from a given sweepout of the sphere by closed curves.
result Lower bounds for the Birkhoff min-max invariant of a Riemannian sphere in terms of the min-max width of its embedded circles.
Improved neural network depth-width trade-offs via dynamical systems.
problem Expressivity of neural networks in terms of depth and width.
method Connection with dynamical systems, focusing on periodic points and Lipschitz constants.
result Sharper width lower bounds for neural networks, yielding exponential depth-width separations.
New findings show depth is more important than width in neural networks.
problem Understanding the role of width and depth in neural networks.
method Constructed networks with bounded weights and width at most d+2, showing depth plays a more significant role.
result Depth is more important than width in the expressive power of neural networks.