Theory proposes neural networks can be initialized for optimal information transmission.
problem Optimizing neural networks for optimal information transmission and representation.
method Developed a corrected mean-field framework to study neural networks as information channels, proving mutual information maximization at dynamic isometry.
result Mutual information maximization is realized between inputs and propagated signals when neural networks are initialized at dynamic isometry.
BraidNet uses braid theory to optimize neural networks for image classification.
problem Image classification problems
method Procedural optimization of neural networks combining information theory and braid theory
result BraidNet outperforms other networks in learning speed and accuracy
IFT reformulates AI and ML tasks using field theory.
problem Signal reconstruction and non-parametric inverse problems.
method Reformulate inference in IFT as GNN training.
result IFT-based GNNs can operate without pre-training.
New method uses information theory to uncover causal relationships in complex systems.
problem Discovering causal relationships in multivariate systems, especially in Bayesian networks and hypergraphs.
method Partial Information Decomposition (PID) to explicitly model higher-order interactions.
result PID components reveal direct causal neighbors and collider relationships in Bayesian networks and multi-tail hyperedges in causal hypergraphs.
New compression theory justifies model pruning for neural networks.
problem Improving neural network performance with reduced model size.
method Information-theoretic rate-distortion theory applied to NN compression.
result Pruning improves model performance on CIFAR-10 and ImageNet datasets.
Recent years, many researches attempt to open the black box of deep neural networks and propose a various of theories to understand it. Among them, Information Bottleneck (IB) theory claims that there are two distinct phases consisting of fitting phase and compression phase in the course of training. This statement att…
A new framework for information theory considers computational constraints.
problem Understanding information in complex systems with computational limitations.
method Variational extension of Shannon's information theory with computational constraints.
result Predictive V-information can be created through computation and reliably estimated from data. Paper introduces a novel error measure for neural networks integrating statistical and information theory.
problem No single error measure is universally best for neural network training.
method Developed a novel error measure EExpAbs and integrated it into the Levenberg-Marquardt algorithm. result Self-adaptive, dynamic learning algorithm improves both model accuracy and training process.
Study on communication delays in decentralized learning networks.
problem Optimizing communication latency in decentralized learning networks.
method Utilized network information theory and random geometric graph theory.
result Communication delay scales as O(n^(2-3β)/βlog n).
The paper connects DNN generalization to node SNR using information theory.
problem Exploring the reasons behind DNN generalization performance.
method Using information theory, the paper derives SNR expressions for DNN nodes and uses them to quantify weight optimization.
result Good SNR performance in DNN nodes correlates with good generalization.
We find the maximum mutual information for neural networks and its key determinants.
problem Understanding the maximum mutual information in neural architectures.
method Derived closed-form expression for maximum mutual information across neural network families.
result Maximum mutual information stems from a generalized formula and is influenced by network width and statistical invariances.
Turbo-Sim generates models from physics principles, improving interpretability and flexibility.
problem Transforming particle properties from theory to observation in collider physics.
method Maximizes mutual information between input and output, setting loss term weights.
result Mathematically interpretable and flexible generative models.
Estimates how much samples inform neural network training and function.
problem Measuring informativeness of samples in neural networks.
method Linearized network approximations for efficient computation.
result Efficient approximations show good accuracy for real-world models.
Two-layer networks learn faster with batch reuse, overcoming information and leap exponents.
problem Limitations of gradient flow and single-pass GD in learning multi-index target functions.
method Multi-pass gradient descent that reuses batches, analyzed using Dynamical Mean-Field Theory.
result Two-time-step overlap with target subspace for non-staircase functions, overcoming information and leap exponents.
Despite great popularity of applying softmax to map the non-normalised outputs of a neural network to a probability distribution over predicting classes, this normalised exponential transformation still seems to be artificial. A theoretic framework that incorporates softmax as an intrinsic component is still lacking. I…
A theoretical framework for deep learning is proposed to explain its effectiveness.
problem Lack of a comprehensive theory explaining deep learning's effectiveness.
method Integrates three characteristics into a graphical model called neurashed.
result Explains common empirical patterns in deep learning and provides insights into regularization and elasticity.
Accurately determining dependency structure is critical to discovering a system's causal organization. We recently showed that the transfer entropy fails in a key aspect of this---measuring information flow---due to its conflation of dyadic and polyadic relationships. We extend this observation to demonstrate that this…
To improve how neural networks function it is crucial to understand their learning process. The information bottleneck theory of deep learning proposes that neural networks achieve good generalization by compressing their representations to disregard information that is not relevant to the task. However, empirical evid…
New bounds improve neural network generalization through slicing.
problem Difficulty in evaluating mutual information in high dimensions for neural networks.
method Slicing the parameter space and using disintegrated mutual information and k-sliced mutual information.
result Slicing improves generalization and offers significant computational and statistical advantages.
Paper studies community detection in censored hypergraphs using information theory.
problem Community detection in censored hypergraphs with missing values.
method Information-theoretic approach, polynomial-time algorithm, spectral algorithm with refinement.
result Derives information-theoretic threshold for exact recovery of community structure.
The Information Plane theory predicts autoencoders do not compress input information.
problem Understanding the training dynamics of hidden layers in autoencoders.
method Derive a theoretical convergence for the Information Plane of autoencoders using a Gram-matrix based mutual information estimator.
result Ideal autoencoders with a large bottleneck layer size do not compress input information, while a small size causes compression only in the encoder layers.
LogDet estimator improves entropy estimation in neural networks.
problem Inconsistent observations and diversified interpretation in neural networks.
method Proposes LogDet estimator for reliable entropy approximation.
result LogDet estimator overcomes distributional diversity issues.
New method estimates spin system mutual information using neural networks.
problem Estimating mutual information in spin systems.
method Monte Carlo sampling enhanced by autoregressive neural networks.
result Area law satisfied for temperatures away from critical temperature.
Review of information plane analyses in neural networks, highlighting mixed results and methodological challenges.
problem Understanding the relationship between information-theoretic compression and neural network performance.
method Literature review and detailed analysis of information quantity estimation methods.
result Information plane compression is not necessarily information-theoretic but compatible with geometric compression.
A measure of neural complexity quantifies how hard it is to access information across neurons.
problem Understanding how mutual information is distributed among neurons in neural networks.
method Partial Information Decomposition (PID) to disentangle contributions of single neurons, multiple neurons, and synergistic effects.
result Representational Complexity measures the difficulty of accessing information across multiple neurons.
Introduces a neural network-based method for efficient state and parameter estimation in complex systems.
problem Efficiently estimating state paths and parameters from noisy measurements in high-dimensional nonlinear systems.
method Bayesian Information Field Theory with neural network parameterization and optimization algorithms.
result Proposes a method to simplify and enrich state path parameterizations using neural networks, improving inference accuracy.
Theory of learning with weight-distribution constraints.
problem Understanding how structure influences function in neural networks.
method Statistical mechanical theory and optimal transport.
result Reduction in capacity due to constrained weight-distribution is related to Wasserstein distance.
Kernel networks' stability edge linked to Fisher Information singularity.
problem Understanding the stability edge in high-capacity kernel Hopfield networks.
method Statistical manifold analysis and Riemannian geometry.
result The Ridge of Optimization corresponds to the Edge of Stability, revealing a dual equilibrium.
Two Fisher information matrix estimators are analyzed for neural networks, focusing on their variances and trade-offs.
problem Estimating the Fisher information matrix in neural networks due to its high computational cost.
method Examined two popular diagonal Fisher information matrix estimators and their variances in neural networks for regression and classification.
result The variances of the estimators depend on the non-linearity with respect to different parameter groups and should not be neglected.
New study shows exponential sample growth for ReQU neural networks.
problem Computing neural network approximations from samples is challenging.
method Information-based complexity tools.
result Functions can be approximated by ReQU neural networks at arbitrary rates but require exponentially growing samples.
Adversarial training enhances model transferability without sacrificing accuracy.
problem The principle of minimal information in classification models is challenged by adversarial training.
method Investigation of the dual relationship between adversarial training and information theory.
result Adversarial training improves linear transferability and introduces a trade-off between transferability and source task accuracy.
Random matrix analysis reveals that neural network weights are mostly random, with some indicating learned information.
problem Understanding how neural networks store information needed for tasks.
method Random matrix theory (RMT) applied to weight matrices of trained deep neural networks.
result Most singular values and eigenvectors of trained neural networks follow universal RMT predictions, suggesting they are random and do not contain system-specific information.
New bounds on neural network convergence using information theory.
problem Quantifying convergence rates of neural networks to Gaussian distributions.
method Entropic inequalities and Gaussian approximations.
result Improved convergence rates in various distances for neural networks.
We study the flow of information and the evolution of internal representations during deep neural network (DNN) training, aiming to demystify the compression aspect of the information bottleneck theory. The theory suggests that DNN training comprises a rapid fitting phase followed by a slower compression phase, in whic…
Networks provide a skeleton for the spread of contagions, like, information, ideas, behaviors and diseases. Many times networks over which contagions diffuse are unobserved and need to be inferred. Here we apply survival theory to develop general additive and multiplicative risk models under which the network inference…
The paper explores how information theory aids in statistical learning models.
problem Characterizing fundamental performance limits in statistical learning models.
method Introduces divergence measures and evidence lower bound (ELBO) in model training.
result Provides a systematic derivation for generative diffusion models.
Study on reducing forgetting in neural networks using compression theory.
problem Catastrophic forgetting in neural networks.
method Defined forgetting as increased description lengths, compared variational posterior approaches to prequential coding methods.
result Proposed a new continual learning method combining ML plug-in and Bayesian mixture codes.
We study the behavior of untrained neural networks whose weights and biases are randomly distributed using mean field theory. We show the existence of depth scales that naturally limit the maximum depth of signal propagation through these random networks. Our main practical result is to show that random networks may be…
Existing information-theoretic frameworks based on maximum entropy network ensembles are not able to explain the emergence of heterogeneity in complex networks. Here, we fill this gap of knowledge by developing a classical framework for networks based on finding an optimal trade-off between the information content of a…
The paper analyzes tensor recovery from symmetric rank-one measurements using information theory.
problem Recovering tensors with low symmetric rank from symmetric rank-one measurements.
method Covering numbers argument, Carbery-Wright inequality, orthogonal polynomials, Fano's inequality.
result Near-optimal sample complexity bounds for log-concave distributions.
While physics conveys knowledge of nature built from an interplay between observations and theory, it has been considered less importantly in deep neural networks. Especially, there are few works leveraging physics behaviors when the knowledge is given less explicitly. In this work, we propose a novel architecture call…
Using established principles from Statistics and Information Theory, we show that invariance to nuisance factors in a deep neural network is equivalent to information minimality of the learned representation, and that stacking layers and injecting noise during training naturally bias the network towards learning invari…
Paper extends information theory for efficient probabilistic modeling.
problem Efficient and data-efficient non-parametric density estimation.
method Structured generative model (SGM) using Rényi's information.
result SGM improves mutual information estimation and generative adversarial networks.
In this work, we investigate the use of three information-theoretic quantities -- entropy, mutual information with the class variable, and a class selectivity measure based on Kullback-Leibler divergence -- to understand and study the behavior of already trained fully-connected feed-forward neural networks. We analyze …
Study reveals how Fisher information changes with network depth, finding it grows linearly.
problem Understanding the trainability of deep neural networks (DNNs).
method Investigates the spectral distribution of the conditional Fisher information matrix (FIM) for fully-connected networks achieving dynamical isometry.
result The conditional FIM's spectrum concentrates around the maximum and grows linearly with depth.
The paper sets information-theoretic lower bounds for neural networks' parameter recovery and excess risk.
problem Establishing sample complexity lower bounds for neural network parameters and excess risk.
method Using information-theoretic tools, the paper proves lower bounds by constructing a generative network.
result Proves information-theoretic lower bounds for exact parameter recovery and positive excess risk.
New approach simulates reputation dynamics using information compression.
problem Malicious communication strategies in reputation networks.
method Uses information compression techniques to simulate social phenomena.
result Emergent phenomena like echo chambers and deception are observed.
The international trade network (ITN) has received renewed multidisciplinary interest due to recent advances in network theory. However, it is still unclear whether a network approach conveys additional, nontrivial information with respect to traditional international-economics analyses that describe world trade only i…