BraidNet uses braid theory to optimize neural networks for image classification.
problem Image classification problems
method Procedural optimization of neural networks combining information theory and braid theory
result BraidNet outperforms other networks in learning speed and accuracy
Braid theory optimizes neural network structures.
problem Optimizing the architecture of neural networks.
method Using braid theory to describe and construct neural network structures.
result Braid-based networks outperform other architectures in classification tasks.
Unified view of spectral networks linking geometry and gauge theory.
problem Understanding BPS states in gauge theories.
method Unified geometric and physical approaches, focusing on spectral networks.
result Spectral networks provide a framework for determining BPS spectra.
Renormalization in neural networks linked to quantum field theory.
problem Implementing renormalization in neural networks.
method Mapping neural networks to quantum field theory, applying renormalization techniques.
result Changing weight standard deviation corresponds to a renormalization flow.
Simplified neural network EFTs reveal a single critical condition.
problem Understanding neuron statistics in neural networks at initialization.
method Diagrammatic approach to effective field theories (EFTs).
result A single condition governs criticality of all neuron preactivations.
The paper examines when NTK theory applies to real finite-width neural networks.
problem Understanding when NTK theory accurately predicts the behavior of finite-width neural networks.
method Empirical study of fully-connected ReLU and sigmoid DNNs with various hyperparameters and depths.
result NTK theory does not always apply to sufficiently deep networks with exploding gradients, and the kernel changes significantly during training.
Paper introduces new neural network models and theories.
problem Understanding neural networks beyond over-parameterized regime.
method Develops two exact models and a novel representor theory.
result Provides insights into neural network training and kernel evolution.
New method uses extreme value theory to estimate neural network errors.
problem Quantifying the error of neural networks, especially for large values.
method Applying extreme value theory to approximate the distribution of error.
result Developed a new estimator for the shape parameter of the Pareto distribution.
This work uses sampling theory to analyze smoothness and error bounds of finite neural networks.
problem Analyzing the function space of finite neural networks and providing error bounds.
method Applying sampling theory to finite neural networks with non-expansive activation functions, considering both deterministic and random sampling.
result Novel error bounds for univariate neural networks under band-limited input assumption, highlighting the advantage of deterministic uniform sampling.
Complex network theory has been applied to solving practical problems from different domains. In this paper, we present a general framework for complex network applications. The keys of a successful application are a thorough understanding of the real system and a correct mapping of complex network theory to practical …
Much attention has been devoted recently to the generalization puzzle in deep learning: large, deep networks can generalize well, but existing theories bounding generalization error are exceedingly loose, and thus cannot explain this striking performance. Furthermore, a major hope is that knowledge may transfer across …
SLT explains neural network success by closing theory-practice gap.
problem Failure of classical inference and learning theory in modern neural networks.
method Physics-inspired Singular Learning Theory (SLT) applied to neural networks.
result SLT recovers known and novel scaling laws for neural network phase transitions.
The abstract proposes a neural network theory using quantum field theory.
problem Understanding the behavior of neural networks in the asymptotic and non-asymptotic limits.
method Mapping neural networks to Wilsonian effective field theory, using Gaussian processes and Feynman diagrams.
result Established a direct connection between overparameterization and simplicity of neural network likelihoods.
Neural nets solve braid untangling up to length 20.
problem Untangling braids in knot theory and group theory.
method Feed-forward neural networks in reinforcement learning.
result Trained neural networks to untangle braids in minimal moves.
New approach connects 3D Chern-Simons theory to spectral networks.
problem Understanding Chern-Simons invariants in 3D manifolds.
method Constructing equivalences between bundles and spectral networks.
result New formulas for Chern-Simons invariants of 3D manifolds.
Quantizes neural networks using frame theory for improved accuracy.
problem Improving neural network efficiency and accuracy through quantization.
method Sigma-Delta (ΣΔ) quantization with finite unit-norm tight frames. result Error bound between original and quantized neural networks derived.
Category theory enhances understanding of group-equivariant neural networks.
problem Understanding and working with group-equivariant neural networks.
method Application of category theory to tensor power spaces of Rn for groups Sn, O(n), Sp(n), and SO(n). result New insights and an algorithm for computing equivariant linear layers.
The paper extends minimal network theory to the sphere, proving local minimality.
problem Finding networks of minimal length on the sphere.
method Adapted spherical geometry, calibration method, and local metric perturbation estimates.
result Spherical minimal networks composed of great-circle arcs are locally length-minimizing within small geodesic balls.
Theory proposes neural networks can be initialized for optimal information transmission.
problem Optimizing neural networks for optimal information transmission and representation.
method Developed a corrected mean-field framework to study neural networks as information channels, proving mutual information maximization at dynamic isometry.
result Mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry.
Deep, wide ConvResNets can approximate functions and their smoothness.
problem Function approximation and smoothness in deep networks.
method Analyzing ConvResNets, proving their ability to approximate functions and their smoothness.
result Large ConvResNets can approximate functions and exhibit sufficient first-order smoothness.
New neural network approach mitigates vanishing/exploding gradients.
problem Vanishing and exploding gradients in neural networks.
method Gaussian-Poincaré normalized functions and orthogonal weight matrices.
result High-dimensional probability theory shows gradients disappear with high probability in wide neural networks.
New method connects curvature and Persistent Homology for networks.
problem Efficient computation of Persistent Homology for complex networks.
method Discrete Morse Theory, Bloch's extension, Forman-Ricci curvature.
result Efficient Persistent Homology scheme using curvature-based approach.
This paper uses neural networks to predict stock prices more accurately.
problem Current stock analysis methods are inaccurate.
method Dynamic neural networks to identify stock price patterns.
result Neural networks outperform traditional stock analysis methods.
Recent years, many researches attempt to open the black box of deep neural networks and propose a various of theories to understand it. Among them, Information Bottleneck (IB) theory claims that there are two distinct phases consisting of fitting phase and compression phase in the course of training. This statement att…
Neural networks outperform kernels by learning features better.
problem Current theories of feature learning do not adequately assess feature quality.
method Introduced feature quality metric and examined existing theories empirically.
result Current theories of feature learning do not provide a sufficient foundation for neural network generalization.
Lecture notes on linear neural networks for deep learning optimization and generalization.
problem Understanding optimization and generalization in deep learning models.
method Mathematical tools and dynamical systems theory.
result Potential of mathematical tools to enhance understanding of deep learning.
Analyzes feature learning in neural networks using a self-consistent dynamical field theory.
problem Feature learning in infinite-width neural networks.
method Constructs deterministic dynamical order parameters as inner-product kernels for hidden unit activations and gradients.
result Reveals the hidden layer activation distribution, neural tangent kernel evolution, and output predictions.
Study shows LLC correlates with neural network compressibility.
problem Evaluating limits of neural network compression.
method Extended minimum description length principle using singular learning theory.
result Complexity estimates based on LLC are linearly correlated with compressibility.
GNNs learn graph representations, with new theory on their power and limitations.
problem Understanding the capabilities and limitations of GNNs.
method Theoretical analysis of GNNs, focusing on approximation and learning properties.
result New insights into the representation, generalization, and extrapolation of GNNs.
We propose using category theory to unify deep learning architectures.
problem Lack of a coherent bridge between model constraints and implementations.
method Apply category theory to unify neural network design.
result Theory recovers constraints from geometric deep learning and encodes standard constructs.
Field theory explains optimal scaling in ResNets for signal propagation.
problem Understanding optimal scaling parameter for ResNet performance.
method Finite-size field theory for ResNets to study signal propagation and scaling.
result Analytical expressions for optimal scaling parameter, independent of other hyperparameters.
Quantum field theory connects deep neural networks to criticality.
problem Understanding the criticality and training dynamics of deep neural networks.
method Constructing quantum field theory for deep neural networks, computing corrections to correlation functions.
result Found precise analogy with O(N) vector model, providing corrections to correlation length. Network theory assesses systemic risk in the insurance sector.
problem Detecting critical insurance companies in systemic risk.
method Complex network approach with weighted effective resistance centrality.
result Identifies companies with significant influence on network robustness.
Fixed points of nonnegative neural networks are analyzed using fixed point theory.
problem Analyzing fixed points in nonnegative neural networks.
method Fixed point theory, nonlinear Perron-Frobenius theory, monotonic and scalable mappings.
result Conditions for the existence of fixed points in nonnegative neural networks are provided.
Deep ReLU networks can approximate matrix-vector products with error bounds.
problem Can deep ReLU networks accurately approximate matrix-vector products?
method Derived error bounds in Lebesgue and Sobolev norms for deep ReLU FNNs.
result Developed deep approximation theory with successful applications.
This paper optimizes sports betting strategies using neural networks and portfolio theory.
problem Optimizing betting strategies in sports gambling.
method Combining neural network models with portfolio optimization, integrating Von Neumann-Morgenstern Expected Utility Theory and the Kelly Criterion.
result Achieved 135.8% relative profit during the English Premier League season.
A new GCN model detects cryptocurrency fraud by considering network evolution and balance theory.
problem Detecting fraud in evolving signed cryptocurrency trust networks.
method Motif-aware temporal GCN using balance theory and learnable weights.
result The model outperforms existing methods on bitcoin datasets.
Neural networks are mathematically represented via quiver representations.
problem Understanding how neural networks process data and create representations.
method Representing neural networks as quiver representations with activation functions.
result Neural networks' computations can be studied algebraically and geometrically.
Random Matrix Theory explains loss surface Hessians in neural networks.
problem Understanding the loss surfaces of neural networks.
method Investigation of local spectral statistics of neural network Hessians.
result Excellent agreement with Gaussian Orthogonal Ensemble statistics.
Equivariant neural networks improve performance and generalization in complex scalar field theory tasks.
problem Improving performance and generalization in neural networks for complex scalar field theory tasks.
method Incorporating translational equivariance into neural network architectures.
result Equivariant neural networks significantly outperform non-equivariant networks in various tasks, including those beyond the training set and across different lattice sizes.
New framework explains deep neural networks using variational spline theory.
problem Understanding functions learned by deep neural networks.
method Developed a variational framework and function space.
result Deep ReLU networks are solutions to regularized data fitting problems over the proposed function space.
Group equivariant neural networks simplify complex tasks with group representation theory.
problem Challenging tasks requiring input transformations like rotations.
method Group representation theory, non-commutative harmonic analysis, differential geometry.
result A neural network is group equivariant if and only if it has a convolutional structure.
Machine learning explores symmetries in field theory and algebra.
problem Understanding symmetries in field theory and algebra.
method Using neural networks to analyze conformal field theory and Lie algebra representation theory.
result Recent advances in machine learning have uncovered new symmetries.
L-CNNs maintain gauge symmetry on non-Abelian lattice theories.
problem Applying convolutional neural networks to non-Abelian lattice gauge theories while preserving gauge symmetry.
method Developed a geometric formulation of L-CNNs that are equivariant under global symmetries and gauge transformations.
result Convolutional operations in L-CNNs are a specific case of gauge-equivariant neural networks on SU(N) principal bundles. In this paper, we theoretically prove that gradient descent can find a global minimum of non-convex optimization of all layers for nonlinear deep neural networks of sizes commonly encountered in practice. The theory developed in this paper only requires the practical degrees of over-parameterization unlike previous the…
Paper characterizes gradient descent dynamics for neural networks with finite width.
problem Characterize gradient descent dynamics for multi-layer neural networks.
method Non-asymptotic state evolution theory for finite-width networks.
result Gradient descent dynamics provide precise distributional characterization.
Equivariant neural networks use symmetry to interpret complex data.
problem Interpreting and understanding the behavior of equivariant neural networks.
method Decompose layers into simple representations and analyze nonlinear activation functions.
result Equivariant neural networks can be interpreted using a filtration generalizing Fourier series.
Theory explains how deep nets learn features from data.
problem Understanding how deep neural networks learn features from data.
method Developed a noise-nonlinearity phase diagram and a mechanical theory.
result Links feature learning across layers to generalization.