Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

4795142189 · Jun 202019922001200920182026
48 results for width constraints

Residual networks with block width max(d_x, d_y) approximate all functions.

problem Achieving universal approximation with residual networks.
method Established bounds on block width for different activation functions.
result Minimum block width for universal approximation is max(d_x, d_y) with inner width 1.

The paper proves a theorem linking ball volume and Uryson width, with applications to 3-sphere metrics.

problem Relating volume of balls to Uryson width and Hausdorff content.
method Short proof and generalization of Guth's theorem, using Riemannian metrics and map diameter constraints.
result For any C>0C>0, there exists a 3-sphere metric with volume 1 and a map violating diameter constraints.

Deep ReLU nets with width d+1 can approximate any convex function on [0,1]^d.

problem Approximating continuous functions on the unit cube with ReLU nets.
method Observing the convexity of ReLU activations and proving approximation by nets of width d+1.
result ReLU nets with width d+1 can approximate any continuous convex function on [0,1]^d arbitrarily well.

Unified spectral framework for μP under joint width-depth scaling.

problem Challenges in stable feature learning and HP transfer for width-depth scaled models.
method Developed a simple and unified spectral framework for μP under joint width-depth scaling.
result Unified and generalized μP formulation for practical architectures with multi-transformation branches.

Joslim optimizes both width and weight configurations for slimmable neural networks, improving model efficiency.

problem Optimizing both width and weight configurations for slimmable neural networks to improve efficiency.
method Proposes a general framework for joint optimization of width configurations and weights, and introduces Joslim algorithm.
result Improves model efficiency by up to 1.7% in top-1 accuracy on the ImageNet dataset.

New metric properties show volume constraints in collapsing spaces.

problem Volume constraints in collapsing spaces.
method Generalization of recent progress in metric geometry involving the volume of balls of radius in a certain range with collapsing at different scales.
result For every Riemannian metric on a manifold of sufficiently small volume, there is a point with volume constraints in the universal cover.

Random groups prove length constraints on product of conjugates.

problem Quantify products of conjugates in random groups.
method Sharp van Kampen diagram argument and boundary block-counting.
result Prove a sharp inequality for products of conjugates in random groups.

Sharp dimension constraints for positive intermediate curvature metrics are established.

problem Proving sharp dimension constraints for metrics with positive intermediate curvature.
method Constructing counterexamples and extending rigidity results.
result Sharp dimension constraints for positive intermediate curvature metrics are established.

Improved sample complexity bounds for neural networks with depth independence.

problem Understanding the sample complexity of neural networks with depth and size independence.
method New bounds on Rademacher complexity with norm constraints on parameter matrices.
result Improved sample complexity bounds that are fully independent of network size under certain assumptions.

Improves matrix multiplication throughput for asymmetric bit-width operands.

problem Matrix multiplications between asymmetric bit-width operands, especially 8- and 4-bit, are not efficiently handled by existing SIMD instructions.
method Proposes a new SIMD matrix multiplication instruction that uses mixed precision on inputs (8- and 4-bit) and accumulates into 16-bit output, improving throughput.
result Offers 2x improvement in throughput compared to existing symmetric-operand-size instructions, with negligible overflow.

A remarkable characteristic of overparameterized deep neural networks (DNNs) is that their accuracy does not degrade when the network's width is increased. Recent evidence suggests that developing compressible representations is key for adjusting the complexity of large networks to the learning task at hand. However, t…

2019-12-10abs ↗pdf ↗

Study on predicting sequences with Gaussian constraints, linking to intrinsic volumes and metric complexity.

problem Predicting sequences almost as well as the best Gaussian distribution with mean in a given subset.
method Expressed minimax regret in terms of intrinsic volumes, established comparison inequality for Wills functional, characterized global covering numbers and local Gaussian widths.
result Sharp estimates on the log-Laplace transform of intrinsic volume sequence for a general nonconvex set.

The width ww of a curve γγ in Euclidean space RnR^n is the infimum of the distances between all pairs of parallel hyperplanes which bound γγ, while its inradius rr is the supremum of the radii of all spheres which are contained in the convex hull of γγ and are disjoint from γγ. We use a mixture of topological and…

2016-05-04abs ↗pdf ↗

Estimates generalization error for two-layer ReLU NNs through minimum norm solutions.

problem Estimating generalization error for two-layer ReLU NNs trained by mean squared error.
method Uses minimum norm solutions and Neural Tangent Kernel (NTK) regime to derive generalization error bounds.
result Derives an a priori generalization error bound for two-layer ReLU NNs without requiring exponentially large number of neurons.

Study min-max theory for hypersurfaces with boundary constraints.

problem Finding minimal hypersurfaces with boundary constraints.
method Schoen-Simon-type regularity result for integral varifolds, proving existence of closed hypersurfaces with specific properties.
result Existence of a closed C1,1C^{1,1} hypersurface with codimension 7\geq 7 singular set in the interior.

Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.

problem Comparing BNNs with different widths due to conflicting model properties and inference intractability.
method Empirical comparison of finite- and infinite-width BNNs, analyzing performance under model mismatch.
result Increasing width can hurt BNN performance when the model is mis-specified, and finite-width BNNs generalize better under model mismatch.

VFlow enhances generative flows by augmenting data dimensions for better expressiveness.

problem Tractable generative flows have limited expressiveness due to fixed intermediate dimensions.
method Augment data with extra dimensions and learn a generative flow for both original and augmented data using variational inference.
result VFlow achieves state-of-the-art performance on CIFAR-10 with improved compactness.

New algorithms reduce dueling bandits' regret with neural networks and efficient exploration.

problem Optimizing dueling bandits with neural networks for better performance.
method Combines shallow exploration strategies with neural networks for utility approximation, using iterative self-improvement and spectral analysis to reduce network width.
result Achieves sublinear regret of O~(dt=1Tσt2+dT)\widetilde{\mathcal{O}}(d\sqrt{\sum_{t=1}^{T} σ_t^2} + \sqrt{dT}).

Lectures on deep learning properties in infinite and large-width networks.

problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.

A number of results for C2^2-smooth surfaces of constant width in Euclidean 3-space E3{\mathbb{E}}^3 are obtained. In particular, an integral inequality for constant width surfaces is established. This is used to prove that the ratio of volume to cubed width of a constant width surface is reduced by shrinking it along…

2007-04-24abs ↗pdf ↗

Gradient descent and SGD achieve low test error in specific network weight regimes.

problem Optimizing two-layer ReLU networks with standard initialization.
method Gradient flow and stochastic gradient descent, analyzing margins and weight norms.
result Gradient descent and SGD can achieve globally maximal margins under certain constraints.

While studying the existence of closed geodesics and minimal hypersurfaces in compact manifolds, the concept of width was introduced in different contexts. Generally, the width is realized by the energy of the closed geodesics or the volume of minimal hypersurfaces, which are found by the Minimax argument. Recently, Ma…

2016-12-20abs ↗pdf ↗

Study on Gaussian-width complexity on statistical manifolds and its applications in learning and recovery.

problem Understanding the geometry of statistical manifolds and its implications for learning and recovery.
method Analysis of Fisher width and inverse-Fisher width, proving their complementary roles and establishing a relation between them.
result Established a sharp relation between Fisher width and inverse-Fisher width, showing they cannot reduce relative to Euclidean scale.

Study infinite-depth limits of neural networks with fixed width.

problem Understanding the behavior of neural networks as depth increases with fixed width.
method Analyzing finite-width residual networks with random Gaussian weights, focusing on the infinite-depth limit.
result The pre-activations converge to a zero-drift diffusion process, differing from the infinite-width limit.

Sparse logistic regression recovers any discrete pairwise graph model.

problem Recovering the Markov graph of discrete pairwise graphical models.
method Maximum conditional log-likelihood with convex optimization.
result The algorithm can recover any arbitrary discrete pairwise graphical model.

In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…

2010-08-30abs ↗pdf ↗

In this paper, we present a unified analysis of matrix completion under general low-dimensional structural constraints induced by {\em any} norm regularization. We consider two estimators for the general problem of structured matrix completion, and provide unified upper bounds on the sample complexity and the estimatio…

2016-03-29abs ↗pdf ↗

Wide neural networks can degrade performance, contrary to conventional wisdom.

problem Understanding the limitations of increasing network width in neural networks.
method Using Deep Gaussian Processes to decouple capacity and width, analyzing their effects on representational power and non-Gaussianity.
result Wide neural networks can become less adaptable and more Gaussian, leading to performance degradation.