Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

8.3%16.7%25.0%33.3% · Jul 199219922001200920172026
48 results for mean width

Bayesian neural networks learn efficiently at infinite width, matching polynomial-width performance.

problem Understanding the inductive bias of infinite-width neural networks.
method Analyzing the reduced entropy and using subsampling techniques.
result The Bayesian mean-field learner generalizes exactly on polynomially-bounded targets.

Study shows bound on Uryson width for specific 3D manifolds.

problem Bounding Uryson width for 3D manifolds with non-negative Ricci curvature and strictly mean convex boundary.
method Proved existence of a Morse function with uniform diameter bounds on level sets.
result Upper bound on Uryson width for the specified 3D manifolds.

Study shows polynomial-width neural networks can closely approximate infinite-width networks in polynomial time.

problem Approximating dynamics of polynomial-width neural networks with infinite-width networks.
method Bounding approximation gap through a differential equation governed by mean-field dynamics, considering local Hessian.
result Polynomially many neurons are sufficient to closely approximate mean-field dynamics.

New framework for understanding infinite-width neural networks.

problem Understanding the infinite-width limit behavior of neural networks.
method General framework to study limit behavior of neural models based on hyperparameter scaling.
result Derives scaling for existing mean-field and neural tangent kernel limits and introduces new dynamically stable limits.

The oloid is the convex hull of two circles with equal radius in perpendicular planes so that the center of each circle lies on the other circle. We calculate the mean width of the oloid in two ways, first via the integral of mean curvature, and then directly. Using this result, the surface area and the volume of the p…

2016-04-25abs ↗pdf ↗

Three configurations of two perpendicular disks in R^3 are examined, the first in which the disks share centers and the other two in which the disks touch at precisely one point. Volume, surface area and mean width calculations dominate the discussion. Integrated mean curvature also appears as an indirect way to comput…

2012-11-19abs ↗pdf ↗

This paper focuses on curves and surfaces of constant width, with some additional results about general ovals. We emphasize the use of Fourier series to derive properties, some of which are known. Amongst other results, we show that the perimeter of an oval is ππ times its average width, and provide a bound for the ra…

2015-04-25abs ↗pdf ↗

We prove the precise scaling, at finite depth and width, for the mean and variance of the neural tangent kernel (NTK) in a randomly initialized ReLU network. The standard deviation is exponential in the ratio of network depth to width. Thus, even in the limit of infinite overparameterization, the NTK is not determinist…

2019-09-13abs ↗pdf ↗

Study on fluctuations in neural network kernels and predictions, focusing on finite width effects.

problem Characterizing fluctuations in finite width neural networks.
method Dynamical mean field theory analysis of wide but finite feature learning neural networks.
result Fluctuations in kernels and predictions are dynamically coupled, leading to reduced variance in feature learning regimes.

The study examines spectral dynamics in deep neural networks, predicting how outliers evolve during training.

problem Understanding spectral evolution in deep neural networks during training.
method Developed a two-level dynamical mean-field theory (DMFT) to track spectral dynamics.
result The theory predicts how outliers evolve with training time, width, output scale, and initialization variance.

Study of manifolds with specific curvature properties using capillary surfaces.

problem Obtaining geometric properties of manifolds with nonnegative scalar curvature and strictly mean convex boundary.
method Use of stable capillary surfaces and Urysohn width to study geometric properties.
result Obtained an obstruction to filling 2-manifolds by 3-manifolds.

We construct a compact, convex ancient solution of mean curvature flow in Rn+1\mathbb R^{n+1} with O(1)×O(n)O(1)\times O(n) symmetry that lies in a slab of width ππ. We provide detailed asymptotics for this solution and show that, up to rigid motions, it is the only compact, convex, O(n)O(n)-invariant ancient solution that lies …

2017-05-19abs ↗pdf ↗

Study shows critical width for rigidity of equatorial zones on spheres.

problem Mean curvature rigidity of equatorial zones on spheres.
method Used tangency principle and trap-slice lemma for strong rigidity, and constructed nontrivial perturbations using Delaunay surfaces for non-rigidity.
result Critical width exists for rigidity, beyond which zones are non-rigid.

Surface area and mean width of a cylinder (the convex hull of two parallel disks) in R^3 are computed. It is more difficult to obtain analogous results for a cone (the convex hull of a disk D and a point p). Oblique formulas for mean width, as well as those for mean curvature, are new. Let L denote the unique diameter …

2012-12-24abs ↗pdf ↗

New framework connects two neural network theories, improving finite-width approximations.

problem Theoretical guarantees for neural network training in general cases.
method Developed a general framework linking mean-field and constant kernel theories.
result Discrete-time MF limit provides better approximation for finite-width nets.

Residual networks with depthwise hyperparameter scaling transfer optimal hyperparameters across width and depth.

problem The challenge of hyperparameter tuning in deep learning, especially for large models.
method Combining μμP parameterization with residual networks having a residual branch scale of 1/extdepth1/\sqrt{ ext{depth}}.
result Optimal hyperparameters transfer across width and depth in residual networks trained with this parameterization.

We give a bound on the extinction time for a compact, strictly convex hypersurface in R^{n+1} evolving by a geometric flow where the velocity is given in terms of the curvature. This result generalizes a theorem of Colding and Minicozzi for mean curvature flow solutions to a wider class of flows studied by Ben Andrews.…

2008-05-07abs ↗pdf ↗

Geodesic balls with non-negative Ricci curvature have a sharp lower bound on their first Dirichlet eigenvalue.

problem Finding a sharp lower bound for the first Dirichlet eigenvalue of geodesic balls.
method Quantitative explicit inequality linking the width of geodesic balls to the spectral gap.
result A quantitative inequality relating the width of geodesic balls to the spectral gap between the first Dirichlet eigenvalue and its lower bound.

The betting CI outperforms classical methods in constructing confidence intervals for bounded means.

problem Constructing nonasymptotic confidence intervals for bounded means.
method A betting-based approach to define and time-uniform variants of confidence intervals (CSs).
result The betting CI matches the fundamental limits, outperforming existing empirical Bernstein CIs.

Study on random linear programs and their connection to mean widths of random polyhedrons.

problem Characterizing the objectives of random linear programs and their relation to mean widths of random polyhedrons.
method Utilizing random duality theory, the exact characterizations of linear objectives are obtained in a large dimensional context.
result The exact characterizations of the program's objectives are obtained, connecting the objectives to the mean widths of random polyhedrons.

This work studies fluctuation in multilayer neural networks using mean field theory.

problem Understanding fluctuation in multilayer neural networks with mean field training.
method Developed a second-order mean field limit to capture fluctuation, demonstrating stability of gradient descent training.
result Gradient descent training in multilayer networks biases towards minimal fluctuation, even after convergence.

Study on neural network dynamics in high dimensions with quadratic activation.

problem Understanding training dynamics in overparameterized neural networks.
method Derivation of gradient flow equations and analysis under l2-regularization.
result Characterization of estimator performance and spectral properties in the high-dimensional limit.

Despite their prevalence in neural networks we still lack a thorough theoretical characterization of ReLU layers. This paper aims to further our understanding of ReLU layers by studying how the activation function ReLU interacts with the linear component of the layer and what role this interaction plays in the success …

2018-12-06abs ↗pdf ↗

Given a Riemannian metric on a homotopy nn-sphere, sweep it out by a continuous one-parameter family of closed curves starting and ending at point curves. Pull the sweepout tight by, in a continuous way, pulling each curve as tight as possible yet preserving the sweepout. We show: Each curve in the tightened sweepout …

2007-05-25abs ↗pdf ↗

New optimizers control network width scaling, improving stability and transfer across different model sizes.

problem Designing stable optimizers for networks of varying widths.
method Interpreting optimizers as steepest descent under mean-normalized operator norms, enabling layerwise composability and width-independent bounds.
result New optimizers like row normalization and column normalization provide stable learning-rate transfer across different model widths.

The study reveals a transition in neural network performance from infinite-width to variance-limited behavior as dataset size increases.

problem Understanding the transition from infinite-width to variance-limited behavior in neural networks.
method Empirical study of the transition from infinite-width to variance-limited behavior as a function of sample size and network width.
result The critical sample size \( P^* \) is approximately \( \sqrt{N} \) for polynomial regression with ReLU networks.

The study analyzes deep linear networks from random initialization, capturing dynamics and hyperparameter effects.

problem Understanding training dynamics in deep linear networks from random initialization.
method Theoretical analysis of gradient descent dynamics in deep linear networks with random initialization and large data.
result Captures the 'wider is better' effect and hyperparameter transfer effects, contrasting with neural-tangent parameterization.

In order to investigate the origin of large price fluctuations, we analyze stock price changes of ten frequently traded NASDAQ stocks in the year 2002. Though the influence of the trading frequency on the aggregate return in a certain time interval is important, it cannot alone explain the heavy tailed distribution of …

2006-06-18abs ↗pdf ↗

Large learning rates work surprisingly well in standard parameterization, contrary to theory.

problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.

Study on curvature bounds for specific hypersurfaces in Anti-de Sitter space.

problem Bounding principal curvatures of constant mean curvature hypersurfaces.
method Generalized convex hull concept and quantitative estimates based on width.
result Explicit bounds on sectional curvature and quasiconformal dilatation.

This paper addresses the so-called conformal capacities in Rn\mathbb R^n, n3n\ge 3, through comparing three existing definitions (due to Betsakos, Colesanti-Cuoghi, Anderson-Vamananmurthy-Fuglede respectively) and studying their associated iso-capacitary inequalities with connection to half-diameter, mean-width, mean-c…

2013-09-14abs ↗pdf ↗

Proves properties of neural network basins of attraction and their expressiveness.

problem Characterize the properties of basins of attraction in neural networks.
method Analyzes width-bounded neural networks, proving properties of basins of attraction.
result Boundedness and path-connectedness of basins of attraction under certain conditions.

Adaptive kernels from neural networks improve model performance.

problem Improving neural network performance through adaptive kernels.
method Deriving adaptive kernels from infinite-width neural networks using feature learning and gradient flow training.
result Adaptive kernels achieve lower test loss compared to traditional kernels.

RAmmStein optimizes liquidity management in AMMs by learning to rebalance efficiently.

problem Optimal control of concentrated liquidity in decentralized exchanges.
method Formulates as an optimal control problem, uses Deep Reinforcement Learning with HJB-QVI.
result Achieves highest net ROI (1.60%) compared to greedy strategies, reduces rebalancing frequency by 85%.

Paper characterizes gradient descent dynamics for neural networks with finite width.

problem Characterize gradient descent dynamics for multi-layer neural networks.
method Non-asymptotic state evolution theory for finite-width networks.
result Gradient descent dynamics provide precise distributional characterization.

Global convergence proved for three-layer neural networks in mean field regime.

problem Optimization efficiency of multilayer neural networks in the mean field regime.
method Developed a rigorous framework for mean field limit of three-layer networks using stochastic gradient descent and neuronal embedding.
result Global convergence guarantee for unregularized feedforward three-layer networks in the mean field regime.

Bayesian linear networks reveal optimal depth and width trade-offs.

problem Understanding how depth, width, and dataset size affect model quality in linear networks.
method Zero noise Bayesian inference with Gaussian weight priors and mean squared error.
result Optimal predictions at infinite depth and maximized Bayesian model evidence at infinite depth.

It has long been known that a single-layer fully-connected neural network with an i.i.d. prior over its parameters is equivalent to a Gaussian process (GP), in the limit of infinite network width. This correspondence enables exact Bayesian inference for infinite width neural networks on regression tasks by means of eva…

2017-11-01abs ↗pdf ↗