Gradient descent with polylogarithmic width achieves arbitrarily low test error for shallow ReLU networks.
problem Achieving low test error with shallow ReLU networks using gradient descent.
method Gradient descent with polylogarithmic width and polylogarithmic number of samples.
result Gradient descent achieves arbitrarily low test error with shallow ReLU networks of polylogarithmic width.
New research shows deep ReLU networks can be learned with polylogarithmic width.
problem Learning deep ReLU networks with limited over-parameterization.
method Using gradient descent, the study establishes learning guarantees for networks with polylogarithmic width.
result Deep ReLU networks can be learned with a polylogarithmic width condition, not just a high degree polynomial.
The paper analyzes GD for KANs, deriving bounds for training, generalization, and privacy.
problem Training dynamics, generalization, and privacy properties of KANs.
method Gradient Descent (GD) analysis for two-layer KANs under logistic loss and NTK-separable assumption.
result Polylogarithmic width suffices for GD to achieve optimization and generalization rates under DP.
Mirzakhani volumes of moduli spaces are polylogarithmic.
problem Understanding the volume of moduli spaces of hyperbolic surfaces.
method Expressed as a sum of polylogarithms evaluated at specific points.
result Mirzakhani volumes are polylogarithmic.
Study higher genus polylogarithms under Riemann surface degenerations.
problem Understanding higher genus polylogarithms under degenerations.
method Investigate the Enriquez connection for polylogarithms and show it becomes a known connection for families of Riemann surfaces.
result Higher genus polylogarithms can be described explicitly as power series in deformation parameters and logarithms of families.
Investigates webs related to cluster algebras and polylogarithms.
problem Understanding webs associated with cluster algebras and polylogarithms.
method Introducing AMP webs and analyzing their properties, proving results and conjectures.
result Many webs associated with polylogarithms and cluster algebras are AMP webs.
Paper solves no-swap regret minimization for combinatorial bandits with polylogarithmic dependence on N.
problem Design efficient no-swap regret algorithms for combinatorial bandits with exponentially large action space.
method Introduces a no-swap-regret learning algorithm with polylogarithmic dependence on N and demonstrates efficient implementation.
result Achieves no-swap regret with polylogarithmic dependence on N, resolving an open problem.
This work proves that large models can be compressed significantly without losing performance.
problem Achieving comparable performance with smaller models and less data.
method Developed a universal compression theory for neural networks and datasets.
result Proved that a generic permutation-invariant function can be compressed into a function of polylogarithmic size with vanishing error.
Quantum machine learning can't achieve polylogarithmic runtimes, even with quantum data access.
problem Bounding the minimum number of samples required for supervised quantum learning.
method Statistical learning theory and quantum machine learning algorithms.
result Quantum machine learning algorithms for supervised learning have at most polynomial speedups over classical algorithms.
We prove optimal subspace embedding conjecture up to sub-polylogarithmic factors.
problem Optimal dimension and sparsity of subspace embeddings.
method Iterative decoupling technique to analyze higher-order trace moment bounds.
result Sub-polylogarithmic factors in dimension and sparsity of subspace embeddings.
Gradient descent on shallow neural networks achieves near-optimal generalization error.
problem Optimizing shallow neural networks with minimal width for generalization and stability.
method Gradient descent in the interpolating regime with minimal width.
result Gradient descent achieves near-optimal generalization error with minimal width.
New algorithm reduces regret in online portfolio and quantum state learning.
problem Efficiently learning portfolios and quantum states online with minimal regret.
method BISONS algorithm for online portfolio selection, SCHRODINGER'S BISONS for quantum states, with polylogarithmic regret.
result First efficient algorithm with polylogarithmic regret for online portfolio selection and quantum states.
New study reveals a polynomial penalty for adapting to unknown margin parameters in batched nonparametric bandits.
problem Adapting to an unknown margin parameter in batched nonparametric bandits.
method Introduces the regret inflation criterion and develops RoBIN algorithm to achieve optimal regret inflation.
result The optimal regret inflation grows polynomially with the horizon T, characterized by a convex optimization problem.
QATS efficiently decodes HMMs with polylogarithmic complexity.
problem Efficiently decoding hidden Markov models from noisy observations.
method Divide-and-conquer procedure with polylogarithmic sequence complexity and cubic state space complexity.
result QATS outperforms Viterbi and PMAP in speed and accuracy.
New method proves fast regret bounds for online RLHF with generalized preferences.
problem Minimizing max-regret in online RLHF with general preferences and bandit feedback.
method Adopted Generalized Bilinear Preference Model (GBPM) to investigate polylogarithmic regret guarantees.
result Proved polylogarithmic regret bounds for Greedy Sampling and Explore-Then-Commit policies under GBPM.
An algorithm calculates Gabai width for thousands of knots.
problem Calculating Gabai width for many knots.
method Algorithmic definition of Wirtinger width leading to efficient Gabai width bounds.
result Proved Wirtinger width equals Gabai width for knots.
Sampling logconcave functions arising in statistics and machine learning has been a subject of intensive study. Recent developments include analyses for Langevin dynamics and Hamiltonian Monte Carlo (HMC). While both approaches have dimension-independent bounds for the underlying continuous processes under s…
In this paper, we study local solutions F=(F1,..,Fn) of a general functional equation of the form F1(U1(x,y))+....+Fn(Un(x,y))=0. A such equation will be called an ``abelian functional equation'' (Afe). We will restrict ourselves to the case when the inner functions Ui's are real rational functions. First we prove that…
Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.
problem Comparing BNNs with different widths due to conflicting model properties and inference intractability.
method Empirical comparison of finite- and infinite-width BNNs, analyzing performance under model mismatch.
result Increasing width can hurt BNN performance when the model is mis-specified, and finite-width BNNs generalize better under model mismatch.
Study shows sample complexity for multicalibration is Θ(ε^-3) with polylogarithmic factors.
problem Minimizing Expected Calibration Error (ECE) for predictors with respect to a family of groups.
method Proved necessary and sufficient sample complexity of Θ(ε^-3) for multicalibration, using online-to-batch reduction and lower bounds.
result Sample complexity of multicalibration is Θ(ε^-3) with polylogarithmic factors, distinguishing it from marginal calibration.
The isospectral problem for p-widths is solved using Zoll metrics on S^2.
problem Determine if a Riemannian manifold is uniquely determined by its p-widths.
method Construct counterexamples on S^2 using Zoll metrics and properties of geodesic p-widths.
result Many counterexamples exist on S^2, showing uniqueness is not guaranteed.
Width trees link link invariants and bridge number.
problem Understanding link invariants through geometric structures.
method Associate width trees to links and use their geometric properties to bound link invariants.
result Width trees uniquely realize certain link invariants under specific conditions.
Lectures on deep learning properties in infinite and large-width networks.
problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.
Computed p-widths for hemisphere, first for manifolds with boundary.
problem Finding p-widths for manifolds with boundary.
method Computed p-widths for the hemisphere.
result First known p-widths for a manifold with boundary.
Polygon p-widths are found via billiard trajectories.
problem Finding p-widths of polygons. method Proved via billiard trajectories and computed specific cases.
result Polygon p-widths are achieved by billiard trajectories. A number of results for C2-smooth surfaces of constant width in Euclidean 3-space E3 are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…
Computed p-widths for real projective plane.
problem Calculating p-widths for real projective plane.
method Standard metric used to compute p-widths.
result Computed p-widths for real projective plane.
Study bounds Urysohn width of manifolds under surgeries.
problem Bounding Urysohn width of manifolds after surgeries.
method Analyzes connected sums and universal covers, applies to general surgeries.
result Optimal constants in estimates of width bounds are shown.
Residual networks with block width max(d_x, d_y) approximate all functions.
problem Achieving universal approximation with residual networks.
method Established bounds on block width for different activation functions.
result Minimum block width for universal approximation is max(d_x, d_y) with inner width 1.
While studying the existence of closed geodesics and minimal hypersurfaces in compact manifolds, the concept of width was introduced in different contexts. Generally, the width is realized by the energy of the closed geodesics or the volume of minimal hypersurfaces, which are found by the Minimax argument. Recently, Ma…
Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.
problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.
Proves conjecture about sphere widths under rotational symmetry.
problem Width stability of rotationally symmetric metrics.
method Proof of conjecture and extensions to higher dimensions.
result Stability of min-max width under rotational symmetry.
New link invariants from diagram colorings match link widths.
problem Defining link widths via diagram colorings.
method Colorings of link diagrams to define invariants and prove their equivalence to link widths.
result Invariants of link widths calculated algorithmically.
Study infinite-depth limits of neural networks with fixed width.
problem Understanding the behavior of neural networks as depth increases with fixed width.
method Analyzing finite-width residual networks with random Gaussian weights, focusing on the infinite-depth limit.
result The pre-activations converge to a zero-drift diffusion process, differing from the infinite-width limit.
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
Sharp lower bound for first Neumann eigenvalue found in terms of diameter and width.
problem Finding the minimum value of the first Neumann eigenvalue for convex domains.
method Proved the sharp lower bound using diameter and width.
result Sharp lower bound for the first Neumann eigenvalue established.
We prove that among all constant width bodies of revolution, the minimum of the ratio of the volume to the cubed width is attained by the constant width body obtained by rotation of the Reuleaux triangle about an axis of symmetry.
Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
Wide neural networks can degrade performance, contrary to conventional wisdom.
problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.
Develops a new theory of width for embedded circles in Riemannian manifolds.
problem Defining and understanding the width of embedded circles in Riemannian manifolds.
method Morse-Lusternik-Schnirelmann theory applied to geodesics and minimising configurations.
result Classifies configurations of minimising geodesics intersecting embedded circles.
Improved GNN simulation of WL test with exponentially lower complexity.
problem Improving the complexity of simulating the Weisfeiler-Lehman test with GNNs.
method Exponentially lower complexity simulation of WL test using GNNs with polylogarithmic parameters and O(log n) bits feature vectors.
result Near-optimal construction with logarithmic lower bounds for feature vector length and neural network size.
Paper proves a noncompact version of Gromov's band-width estimate.
problem Proving a precise upper bound for noncompact Riemannian bands.
method Developed a quantitative partitioned manifold index theory.
result Proved a version of Gromov's band-width estimate for noncompact Riemannian bands.
We discuss a possible definition for "k-width" of both a closed d-manifold Md, and on embedding Md↪eRn, n>d≥k, generalizing the classical notion of width of a knot. We show that for every 3-manifold 2-width(M3)≤2 but that there are embeddings $e_i: T^3 \hoo…
Kähler information manifolds for signal filters in weighted Hardy spaces are explored.
problem Developing a geometric framework for signal processing filters in weighted Hardy spaces.
method Introducing weighted Hardy spaces and smooth transformations of transfer functions, demonstrating the Kähler manifold structure.
result The Riemannian geometry of weighted Hardy norms for transfer functions forms a Kähler manifold.
New algorithms minimize regret in both adversarial and stochastic contexts.
problem Minimizing regret in linear contextual bandits.
method Best-of-both-worlds algorithms using FTRL with Shannon entropy regularizer.
result Achieves near-optimal regret bounds in both adversarial and stochastic regimes.
We extend the classical definition of {\it width} to higher dimensional, smooth codimension 2 knots and show in each dimension there are knots of arbitrarily large width.
Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.
problem Understanding the optimal balance between depth and width in self-attention models.
method Theoretical predictions and empirical ablations on networks of varying depths and widths.
result An optimal width of 30K is recommended for a 1-Trillion parameter network, marking a significant width for self-attention models.
We present an alternative proof of the following fact: the hyperspace of compact closed subsets of constant width in Rn is a contractible Hilbert cube manifold. The proof also works for certain subspaces of compact convex sets of constant width as well as for the pairs of compact convex sets of constant rela…