A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We give upper bounds on the principal curvatures of a maximal surface of nonpositive curvature in three-dimensional Anti-de Sitter space, which only depend on the width of the convex hull of the surface. Moreover, given a quasisymmetric homeomorphism φ, we study the relation between the width of the convex hull of th…
Maximal initial learning rate for deep ReLU networks identified.
problem Finding the optimal initial learning rate for deep neural networks.
method Simple approach to estimate maximal initial learning rate η∗, analyzing its behavior in constant-width fully-connected ReLU networks.
result Maximal initial learning rate η∗ is well predicted as a power of depth × width, with specific conditions for network width and input layer training.
New scaling framework for MoE architectures ensures stability and optimal performance at scale.
problem Lack of principled understanding of how hyperparameters should scale in MoE architectures.
method Developed a novel Dynamical Mean Field Theory (DMFT) for three scaling regimes of MoE architectures.
result Derived Maximally Scale-Stable Parameterization (MSSP) for SGD and Adam, providing robust learning rate transfer and monotonic improvement with scale.
We prove that Ricci flows with almost maximal extinction time must be nearly round, provided that they have positive isotropic curvature when crossed with R2. As an application, we show that positively curved metrics on S3 and RP3 with almost maximal width must be nearly round.
New framework explains fast transfer of hyperparameters across model scales.
problem Understanding and optimizing hyperparameters for large-scale models.
method Developed a conceptual framework for HP transfer across scale, showing fast transfer is equivalent to useful transfer for compute-optimal grid search.
result Fast transfer of hyperparameters is equivalent to useful transfer for compute-optimal grid search, offering asymptotic computational advantage.
We trace the initiative by Professor Meghnad Saha to develop a (statistical) physics model of market economy and his search for the mechanism to constrain the entropy maximized width of the income distribution in a society such that the spread of inequality can be minimized.
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
problem In modern scale-invariant architectures, training quickly enters a steady state where normalization layers create backward scale sensitivity, degrading learning-rate transfer.
method We introduce a weight-decay scaling rule for AdamW that preserves sublayer gain across widths by equalizing the effective learning rate.
result Our empirical weight-decay scaling rule λ2∝d approximately keeps sublayer gains width invariant, enabling zero-shot transfer of learning rate and weight decay.
Semi-supervised clustering aims to introduce prior knowledge in the decision process of a clustering algorithm. In this paper, we propose a novel semi-supervised clustering algorithm based on the information-maximization principle. The proposed method is an extension of a previous unsupervised information-maximization …
Using the parameterisation of the deformation space of GHMC anti-de Sitter structures on S×R by the cotangent bundle of the Teichmüller space of S, we study how some geometric quantities, such as the Lorentzian Hausdorff dimension of the limit set, the width of the convex core and the Hölder exponen…
This paper shows how deep neural networks can learn rich, independent features that significantly deviate from initialization.
problem Understanding how deep neural networks achieve meaningful feature learning and global convergence.
method Investigation of infinitely wide, L-layer neural networks using the tensor program framework under Maximal Update parametrization.
result SGD enables these networks to learn linearly independent features that substantially deviate from their initial values, capturing relevant data information.
We define the Wirtinger width of a knot. Then we prove the Wirtinger width of a knot equals its Gabai width. The algorithmic nature of the Wirtinger width leads to an efficient technique for establishing upper bounds on Gabai width. As an application, we use this technique to calculate the Gabai width of approximately …
Lectures on deep learning properties in infinite and large-width networks.
problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.
A number of results for C2-smooth surfaces of constant width in Euclidean 3-space E3 are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…
While studying the existence of closed geodesics and minimal hypersurfaces in compact manifolds, the concept of width was introduced in different contexts. Generally, the width is realized by the energy of the closed geodesics or the volume of minimal hypersurfaces, which are found by the Minimax argument. Recently, Ma…
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
We prove that among all constant width bodies of revolution, the minimum of the ratio of the volume to the cubed width is attained by the constant width body obtained by rotation of the Reuleaux triangle about an axis of symmetry.
We discuss a possible definition for "k-width" of both a closed d-manifold Md, and on embedding Md↪eRn, n>d≥k, generalizing the classical notion of width of a knot. We show that for every 3-manifold 2-width(M3)≤2 but that there are embeddings $e_i: T^3 \hoo…
We extend the classical definition of {\it width} to higher dimensional, smooth codimension 2 knots and show in each dimension there are knots of arbitrarily large width.