RAF model explains neural networks' dual rule learning and fact memorization.
problem Understanding how neural networks learn rules and memorize facts simultaneously.
method Introduces the Rules-and-Facts (RAF) model to bridge generalization and memorization.
result Characterizes conditions for simultaneous rule learning and fact memorization in neural networks.
We study finite sample expressivity, i.e., memorization power of ReLU networks. Recent results require N N N hidden nodes to memorize/interpolate arbitrary N N N data points. In contrast, by exploiting depth, we show that 3-layer ReLU networks with Ω ( N ) Ω(\sqrt{N}) Ω ( N ) hidden nodes can perfectly memorize most datasets with N N N po…
New measure FTC quantifies how much a ReLU network can fine-tune.
problem Analyzing memorization capacity in fine-tuned neural networks.
method Defined Fine-Tuning Capacity (FTC) for additive fine-tuning of ReLU networks.
result Upper and lower bounds on FTC for 2 and 3-layer ReLU networks.
Derives an empirical capacity model for self-attention neural networks.
problem Theoretical capacity of large transformer models is not fully utilized by current optimization algorithms.
method Analyzes memory capacity of transformers using synthetic training data and common training algorithms.
result Derives an empirical capacity model (ECM) for a generic transformer.
Data selection boosts fact memorization in language models.
problem Language models struggle to accurately memorize factual knowledge.
method Formalizes fact memorization, proposes data selection schemes based on training loss.
result Data selection boosts fact accuracy to model capacity and improves performance.
This paper explores memorization in adversarial training and proposes a mitigation algorithm.
problem Understanding and mitigating robust overfitting in adversarial training.
method Demonstrated the capacity of deep networks to memorize adversarial examples, analyzed convergence and generalization issues, and proposed a new mitigation algorithm.
result Identified robust overfitting as a significant drawback of adversarial training and proposed a mitigation algorithm.
We examine the role of memorization in deep learning, drawing connections to capacity, generalization, and adversarial robustness. While deep networks are capable of memorizing noise data, our results suggest that they tend to prioritize learning simple patterns first. In our experiments, we expose qualitative differen…
New algorithm interpolates data with neural nets, independent of sample size.
problem Understanding neural networks' ability to memorize training data.
method Randomized algorithm for constructing interpolating neural networks.
result Guarantees that are independent of the number of samples, moving beyond worst-case memorization capacity bounds.
We improve deep threshold networks' memorization capacity exponentially.
problem Memorizing datasets with randomized labels using deep neural networks.
method Using Gaussian random weights in the first layer and binary or integer weights in subsequent layers, we prove a new dependence on minimum distance.
result We show that O ~ ( 1 δ + n ) \widetilde{\mathcal{O}}(\frac{1}{\delta} + \sqrt{n}) O ( δ 1 + n ) neurons and O ~ ( d δ + n ) \widetilde{\mathcal{O}}(\frac{d}{\delta} + n) O ( δ d + n ) weights are sufficient. Logarithmic network width suffices for robust memorization.
problem Achieving robust memorization in neural networks.
method Established upper and lower bounds on robust memorization radius.
result Width logarithmic in the number of samples is necessary and sufficient for robust memorization.
Generative diffusion models gradually memorize training data, losing independent dimensions.
problem Understanding how generative diffusion models memorize training data, especially on low-dimensional manifolds.
method Measuring latent dimensionality via the learned score field, proposing a geometric memorization theory.
result Generative diffusion models experience a smooth collapse of their capacity to vary across independent directions as data become scarce, leading to near point-wise replication of salient features.
New analysis tightens memory capacity of Hopfield models using spherical codes.
problem Optimizing memory capacity in modern Hopfield models and Kernelized Hopfield Models.
method Connecting Hopfield models to spherical codes in information theory, establishing an optimal capacity bound and a sub-linear algorithm.
result First tight and optimal asymptotic memory capacity for modern Hopfield models, matching known lower bounds.
Algorithm performance in supervised learning is a combination of memorization, generalization, and luck. By estimating how much information an algorithm can memorize from a dataset, we can set a lower bound on the amount of performance due to other factors such as generalization and luck. With this goal in mind, we int…
Researchers analyze backdoor data poisoning attacks and identify a memorization capacity parameter.
problem Understanding and mitigating backdoor data poisoning attacks in machine learning models.
method Formal theoretical framework, statistical and computational analysis, explicit constructions, and algorithm design.
result Identified a memorization capacity parameter to assess vulnerability to backdoor attacks and developed algorithms to detect and mitigate them.
Improved neural network capacity analysis using simplified RDT.
problem Analyzing the memorization capabilities of sign perceptron neural networks.
method Developed a simplified, partially lifted Random Duality Theory (fl RDT) approach.
result Concrete capacity bounds universally improve over previous best known ones.
New NTK bounds show deep networks with minimum over-parameterization can still memorize and optimize.
problem Understanding memorization and optimization in sub-linear over-parameterized deep networks.
method Lower bound on NTK eigenvalues for deep networks with minimum over-parameterization.
result Deep networks with minimum over-parameterization can still be powerful memorizers and optimizers.
Transformers with CoT don't enhance reasoning power across all tasks.
problem Does CoT enhance the reasoning power of transformers?
method Examined the memorization capabilities of fixed-precision transformers with and without CoT.
result Transformers with CoT cannot memorize all reasoning tasks, leading to a negative answer.
DMs emerge from DenseAMs, transitioning from memorization to generalization.
problem Hindered memory retrieval in DenseAMs due to spurious states.
method Examined diffusion models through the lens of DenseAMs, focusing on their generative process.
result Identified a critical phase in DMs transitioning from memorization to generalization.
A Gaussian mixture model improves generalization for long-tailed data.
problem Optimizing generalization for rare data in long-tailed distributions.
method Suggested Gaussian mixture model and comparison of linear vs. nonlinear classifiers.
result Nonlinear classifiers outperform linear ones for long-tailed data.
Deep neural networks can generalize by reducing high-frequency noise over time, not always following a monotonic learning bias.
problem Understanding the learning dynamics and generalization of over-parameterized DNNs.
method Experimental analysis of deep double descent, focusing on the spectral bias of DNNs.
result The high-frequency components of DNNs diminish over training, leading to a second descent in test error.
Standard Transformers approximate Hölder functions and achieve optimal nonparametric regression rate.
problem Approximating Hölder functions and achieving optimal nonparametric regression rate with Transformers.
method Using the size tuple and dimension vector metrics, the paper characterizes Transformer structures and derives upper bounds for their Lipschitz constant and memorization capacity.
result Standard Transformers achieve the minimax optimal rate in nonparametric regression for Hölder target functions.
Deep learning with noisy labels is practically challenging, as the capacity of deep models is so high that they can totally memorize these noisy labels sooner or later during training. Nonetheless, recent studies on the memorization effects of deep neural networks show that they would first memorize training data of cl…
Overwhelming theoretical and empirical evidence shows that mildly overparametrized neural networks -- those with more connections than the size of the training data -- are often able to memorize the training data with 100 % 100\% 100% accuracy. This was rigorously proved for networks with sigmoid activation functions and, very …
Fine-tuning LLMs on privacy-sensitive data introduces privacy risk, and synthetic data audits can quantify this risk.
problem Fine-tuning LLMs on privacy-sensitive data introduces privacy risk.
method Generate synthetic canaries via high-temperature sampling from LLMs.
result Synthetic canaries are high-influence outliers that ensure strong audits.
Transformers mimic Bayesian reasoning in controlled settings, revealing geometric mechanisms.
problem Verifying if transformers perform Bayesian reasoning rigorously in natural data.
method Constructing Bayesian wind tunnels with known posteriors and proving memorization impossibility.
result Transformers achieve 10 − 3 10^{-3} 1 0 − 3 - 10 − 4 10^{-4} 1 0 − 4 bit accuracy in Bayesian posteriors, while MLPs fail. New bounds on neural network capacity for treelike sign perceptrons using RDT.
problem Determining the capacity of treelike sign perceptrons neural networks.
method Random Duality Theory (RDT) to establish upper bounds.
result Mathematically rigorous bounds on network capacity for any number of neurons.
Neural networks memorize exceptions, leading to poor generalization.
problem Memorization of exceptions hinders neural network generalization.
method Formalized memorization-generalization interplay, proposed MAT to shift logits.
result MAT improves generalization by learning robust patterns invariant across distributions.
Local data coverage governs memorization in diffusion models.
problem Memorization in diffusion models
method Derive a theoretical criterion based on local data coverage
result Predicts memorization based on density of training data in neighborhood and dataset size
Geometric framework explains memorization in generative models.
problem Memorization in generative models raises legal and privacy concerns.
method Manifold memorization hypothesis (MMH) using manifold geometry.
result Formalizes and categorizes memorization into overfitting and distribution-driven types.
The paper identifies patterns in language model weights used for memorizing paragraphs.
problem Locating the specific mechanisms and weights used by language models to memorize paragraphs.
method Examined gradients and attention patterns in language models to identify memorized paragraphs.
result Gradients of memorized paragraphs have a distinguishable spatial pattern, and localized attention heads are involved in paragraph memorization.
Consistency distillation reduces memorization in diffusion models without harming sample quality.
problem Understanding how distillation affects memorization in diffusion models.
method Analysis of consistency distillation in diffusion models using a random feature neural network model.
result Consistency distillation reduces memorization in diffusion models without harming sample quality.
Study shows how deep generative models can memorize data.
problem Understanding and preventing memorization in deep generative models.
method Adapted a memorization measure for unsupervised density estimation and demonstrated its effectiveness.
result Memorization in deep generative models differs from mode collapse and overfitting.
Learning to remember long sequences remains a challenging task for recurrent neural networks. Register memory and attention mechanisms were both proposed to resolve the issue with either high computational cost to retain memory differentiability, or by discounting the RNN representation learning towards encoding shorte…
Recently deep neural networks have shown their capacity to memorize training data, even with noisy labels, which hurts generalization performance. To mitigate this issue, we provide a simple but effective baseline method that is robust to noisy labels, even with severe noise. Our objective involves a variance regulariz…
Memorizing rare examples helps neural networks generalize better.
problem Improving generalization in deep learning models.
method Theoretical analysis and experiments on neural networks with composition capability.
result Memorizing rare examples can help neural networks make correct predictions on rare test examples.
Optimal ReLU networks can memorize any separable set of points with a small number of parameters.
problem The optimal number of parameters required to memorize a set of points using ReLU networks.
method Construction of ReLU networks with specific bit complexity to memorize points satisfying a mild separability assumption.
result Optimal ReLU networks can memorize any separable set of points with a number of parameters that is i l d e O ( N ) ilde{O}(\sqrt{N}) i l d e O ( N ) . LLMs can memorize economic data and recall exact values before their training cutoff.
problem Evaluating the trustworthiness of LLMs' economic forecasts during their training period.
method Demonstrated through counterfactual forecasting and analysis of LLMs' recall ability.
result LLMs have memorized economic and financial data, leading to recall-level accuracy before their knowledge cutoff.
New method reduces memorization in diffusion models without sacrificing image quality.
problem Diffusion models often memorize training data, especially with small datasets.
method Train models using noisy data at large noise scales to reduce memorization.
result Significant reduction in memorization without compromising image quality.
Paper explains adversarial training's robust overfitting through a minimax game perspective.
problem Adversarial training suffers from robust overfitting after learning rate decay.
method Viewing adversarial training as a dynamic minimax game, analyzing how LR decay breaks balance and leads to overfitting.
result ReBalanced Adversarial Training (ReBAT) alleviates robust overfitting without sacrificing robustness.
New approach shows data memorization trade-offs in large models.
problem Data memorization in large language models and its privacy implications.
method Developed a new approach using strong data processing inequalities to prove lower bounds on memorization.
result Proved that Ω ( d ) Ω(d) Ω ( d ) bits of training data information must be memorized for O ( 1 ) O(1) O ( 1 ) examples, decaying with example growth. Study compares memorization of SimCLR to supervised and random labels training.
problem Understanding memorization in contrastive learning.
method Investigated SimCLR's memorization properties compared to supervised and random labels training.
result SimCLR's memorization is similar to random labels training in terms of training object complexity distribution.
Deep networks preferentially learn shared features, avoiding memorization in early layers.
problem Understanding how deep neural networks generalize vs. memorize training data.
method Replica-based mean field geometric analysis of deep neural networks.
result Deep layers predominantly memorize, while early layers are minimally affected.
Deep networks can memorize random labels; symmetric loss improves this.
problem Deep networks can memorize random labels, ignoring standard regularization.
method Empirical studies with MNIST and CIFAR-10 datasets, formal definition of robustness.
result Symmetric loss function improves network's ability to resist memorization.
Study shows FL reduces unintended memorization by clustering data and using strong user-level privacy.
problem Unintended memorization in federated learning.
method Examined the effect of clustering data and using strong user-level differential privacy in FL.
result Clustering data and strong user-level differential privacy reduce unintended memorization.
A new model improves recurrent neural networks' ability to memorize long sequences.
problem Improving recurrent neural networks' ability to memorize long sequences and extract task-relevant features.
method Proposes a Linear Memory Network with an encoding-based memorization component and a specialized training algorithm.
result Improves the final performance of recurrent neural networks when memorizing long sequences is necessary.
ResMem improves model generalization by explicitly memorizing residuals.
problem Improving model generalization in neural networks.
method ResMem algorithm that augments a model with a k-nearest neighbor based regressor fitted to residuals.
result ResMem consistently improves test set generalization across various benchmarks.
Diffusion models can memorize training data, limiting their creativity and privacy.
problem Memorization in diffusion models that reproduces training data instead of generating novel outputs.
method Dual-separation approach via statistical estimation and network approximation.
result Pruning-based method reduces memorization while maintaining generation quality.
Gradient descent memorizes many Gaussians efficiently.
problem Memorizing many Gaussians with minimal parameters.
method Gradient descent on a depth-two neural network.
result One step of gradient descent memorizes $Ω\left(\frac{dq}{\log^4(d)}
ight)$ Gaussians.