Pruning neural networks can improve test accuracy even with significant parameter reduction.
problem The tradeoff between generalization and stability in neural network pruning.
method Analysis of pruning behavior over training, focusing on instability and its relation to generalization.
result Pruning's benefit to generalization increases with its instability.
A new method to prune neural networks with iterative randomization improves efficiency.
problem Efficiency in pruning randomly initialized neural networks.
method Iteratively randomizing weight values to reduce parameter requirements.
result The method achieves remarkable performance with fewer parameters.
A new framework explains why early pruning works well.
problem Understanding why early pruning of neural networks leads to good performance.
method Gradient flow framework to unify pruning measures.
result Magnitude-based pruning removes least contributing parameters, leading to faster convergence.
Pruning neural networks reduces parameters without sacrificing interpretability.
problem Reducing unnecessary structure in neural networks to improve efficiency.
method Examined the effect of pruning on the number of hidden units learning disentangled representations.
result Pruning does not harm interpretability until a significant portion of parameters are removed.
ICE-Pruning accelerates deep neural network pruning by 9.61x.
problem Efficiently pruning deep neural networks while maintaining accuracy.
method Iterative pruning with automatic fine-tuning steps, freezing strategy, and custom learning rate scheduler.
result Significantly reduces pruning time by up to 9.61x.
New statistical mechanics analysis shows edge pruning outperforms node pruning in neural networks.
problem Theoretical understanding of neural network pruning effectiveness is lacking.
method Statistical mechanics analysis of a teacher-student framework.
result DPP node pruning method is superior to other methods, but edge pruning is better overall.
This work proposes efficient unit pruning for neural networks to reduce model size.
problem Redundant parameters in neural networks can be pruned without performance loss.
method Proposes unit-wise pruning over parameter pruning, introduces saliency scores, and defines dead units.
result 5x model size reduction on MNIST using unit-wise pruning.
A model predicts how hyperparameters affect pruning performance.
problem Predicting the impact of hyperparameters on pruning performance.
method Phenomenological model using temperature-like and load-like parameters.
result A sharp transition phenomenon in pruning performance.
Alpha-trimming prunes trees in random forests to improve predictive performance.
problem Improving predictive performance of random forests by locally adaptive tree pruning.
method Alpha-trimming is a fast pruning algorithm that prunes trees in a random forest based on signal-to-noise ratio, controlled by a tuning parameter.
result Alpha-trimming often lowers mean squared prediction error compared to fully grown random forests.
Artificial neural networks (ANNs) may not be worth their computational/memory costs when used in mobile phones or embedded devices. Parameter-pruning algorithms combat these costs, with some algorithms capable of removing over 90% of an ANN's weights without harming the ANN's performance. Removing weights from an ANN i…
ANPyC combats forgetting by pruning and consolidating neural parameters.
problem Catastrophic forgetting in neural networks, especially with long-term tasks.
method Adversarial Neural Pruning and Synaptic Consolidation (ANPyC) to balance task-relevant and irrelevant parameters.
result ANPyC prevents forgetting while enabling efficient learning of multiple tasks.
New insights on pruning deep networks by preserving function locality.
problem Designing effective pruning methods for deep neural networks.
method Revisited loss modeling using first and second order Taylor expansions, emphasizing locality.
result Both first and second order Taylor expansions can achieve similar performance in pruning.
New pruning method retains model expressiveness for NLP tasks.
problem Pruning large pretrained transformer models for real-world deployment.
method Mixture Gaussian Prior Pruning (MGPP) algorithm.
result MGPP outperforms existing pruning methods in high sparsity settings.
Renormalized pruning improves neural network accuracy.
problem Over-parameterized neural networks waste many parameters.
method Propose renormalizing sparse neural networks.
result Renormalized pruning converges to zero error.
Global compression method prunes deep neural networks efficiently.
problem Efficiently compress deep neural networks for resource-constrained devices.
method Global Sparse Momentum SGD for on-the-fly pruning.
result Automatically finds appropriate per-layer sparsity ratios without re-training.
TENP prunes experts and neurons in Mixture-of-Experts models for efficient deployment.
problem Efficient deployment of large language models constrained by static parameter footprint.
method Structured Trapezoidal ExpertNeuron Pruning (TENP) identifies and retains important experts and neurons.
result DeepSeek model achieves 10% better performance on code generation tasks with 40% expert sparsity.
i-SpaSP prunes neural networks by identifying important groups of parameters, improving pruning efficiency.
problem Pruning neural networks to reduce computational cost and improve performance.
method i-SpaSP uses sparse signal recovery principles to iteratively identify and threshold important parameter groups.
result i-SpaSP achieves strong empirical results and theoretical convergence guarantees, improving pruning efficiency.
A method to prune 3D CNNs by assigning different regularization parameters to layers based on importance.
problem Massive computation and storage consumption in 3D CNNs.
method Regularization-based pruning method assigning different regularization parameters to different weight groups.
result Pruning leads to 2x speedup with minimal accuracy loss for 3DResNet18 and C3D.
Parameter pruning is a promising approach for CNN compression and acceleration by eliminating redundant model parameters with tolerable performance degrade. Despite its effectiveness, existing regularization-based parameter pruning methods usually drive weights towards zero with large and constant regularization factor…
Pruning at initialization fails to find sparse subnetworks, revealing information-theoretic barriers.
problem Difficulty in finding sparse subnetworks without training the full model.
method Analysis of effective parameter count and mutual information between sparsity mask and data.
result Pruning at initialization cannot find sparse subnetworks due to high mutual information.
Prunes neural networks while preserving accuracy, using sensitivity sampling.
problem Sparsifying neural networks while maintaining predictive accuracy.
method Uses sensitivity sampling to construct an importance distribution, then adaptively prunes weights.
result Pruned networks incur minimal loss in performance compared to original networks.
Paper proposes a faster DNN pruning method.
problem Efficiently prune DNNs without sacrificing accuracy.
method Uses PCA to automatically optimize pruning.
result Prunes DNNs in one shot, no retraining needed.
A new energy-efficient pruning method for federated learning.
problem Energy inefficiency in gradient sparsification for federated learning.
method Formalized energy-constrained projection problem and proposed Cost-Weighted Magnitude Pruning (CWMP).
result CWMP optimally balances performance and energy efficiency in federated learning.
Paper proposes efficient network pruning method for deep neural networks.
problem High computational and memory cost of deep neural networks.
method Annealing and direct sparsity control for channel-level pruning.
result Proposed method achieves better or competitive performance compared to other methods.
A taxonomy of saliency metrics helps in pruning deep neural networks.
problem Difficulty in separating saliency metric effectiveness from pruning algorithms.
method Proposed a taxonomy based on four orthogonal components.
result Constructed metrics can outperform existing state-of-the-art metrics.
A method estimates and prunes neural network filters to reduce computation and improve accuracy.
problem Reduction of neural network parameters to save computation and energy.
method Estimates each neuron's contribution to loss using first and second-order Taylor expansions; iteratively removes less important neurons.
result High (>93%) correlation between estimated and true importance; 40% FLOPS reduction with 0.02% top-1 accuracy loss.
ResRep prunes CNNs without losing accuracy by separating remembering and forgetting.
problem Pruning CNNs to reduce FLOPs without sacrificing accuracy.
method Decoupling remembering and forgetting in CNNs, using SGD for remembering and a novel update rule for forgetting.
result Achieved lossless pruning with high compression ratio (76.15% accuracy on ImageNet with 45% FLOPs reduction).
Dirichlet pruning compresses neural networks by removing unimportant units.
problem Compressing large neural network models without sacrificing performance.
method Assigns Dirichlet distribution over network layers' units and uses variational inference to estimate parameters.
result Achieves state-of-the-art compression performance on larger architectures like VGG and ResNet.
This paper studies a theoretical pruning method for RNNs to reduce computational costs.
problem High computational costs in recurrent neural networks (RNNs).
method Spectral pruning inspired approach for RNNs.
result Generalization error bounds for compressed RNNs are provided.
Neural network pruning lacks standardized benchmarks and metrics.
problem Lack of standardized benchmarks and metrics in neural network pruning.
method Meta-analysis of 81 papers, controlled conditions, ShrinkBench framework.
result Neural network pruning community lacks standardized benchmarks and metrics.
New method prunes neural networks at initialization, improving performance.
problem Improving neural network compression at initialization.
method Formally characterizes initialization conditions for reliable pruning based on connection sensitivity.
result Improved neural network performance on image classification tasks.
Synaptic pruning reduces CNNs by 96% on CIFAR-10.
problem Memory and computation constraints in CNNs for mobile devices.
method Synaptic Pruning: data-driven method to prune connections based on Synaptic Strength.
result Significant size reduction and computation saving with up to 96% pruning on CIFAR-10.
In this paper, we propose a novel progressive parameter pruning method for Convolutional Neural Network acceleration, named Structured Probabilistic Pruning (SPP), which effectively prunes weights of convolutional layers in a probabilistic manner. Unlike existing deterministic pruning approaches, where unimportant weig…
Recent DNN pruning algorithms have succeeded in reducing the number of parameters in fully connected layers, often with little or no drop in classification accuracy. However, most of the existing pruning schemes either have to be applied during training or require a costly retraining procedure after pruning to regain c…
Real-time pruning during training reduces network size and training time.
problem Efficiently reducing neural network size without sacrificing accuracy.
method Activation density-based pruning during training.
result Up to 200x reduction in parameters and 60x reduction in inference compute operations.
Pruning neural networks adds differential privacy noise, preserving data utility.
problem Achieving differential privacy in neural networks without sacrificing data utility.
method Proving equivalence between pruning and adding differential privacy noise to hidden-layer activations.
result Pruning can be a more effective alternative to adding differential privacy noise for neural networks.
New pruning technique reduces index size for DNNs.
problem Irregular index form in fine-grained pruning limits parallelism and memory usage.
method Proposes a low-rank binary index matrix for efficient compression and decompression.
result Fine-grained pruning with binary matrices achieves lower memory footprint and higher parallelism.
Hard thresholding remains efficient for DNN pruning, but smart pruning offers faster accuracy recovery.
problem Efficiently pruning deep neural networks while minimizing accuracy loss.
method Proposes a novel smart pruning algorithm based on difference of convex functions optimization.
result Smart pruning is often orders of magnitude faster than competing approaches while achieving low accuracy degradation.
Paper proposes efficient pruning method for neural networks.
problem Compressing deep neural networks for resource-constrained devices.
method Adaptive sparsity loss for budget-aware optimization during training.
result Demonstrated effectiveness on various architectures and datasets.
Paper explores pruning and quantisation to compress neural networks.
problem Reduces computational and memory costs of deep neural networks.
method Investigates network pruning and quantisation for AlexNet, ShuffleNet, and MobileNet.
result Pruning and quantisation compress networks to less than half their size and improve efficiency.
AlphaPruning optimizes LLM pruning using HT-SR theory for better performance.
problem Improving pruning of large language models to reduce size without sacrificing performance.
method AlphaPruning uses HT-SR theory to allocate layerwise sparsity ratios more theoretically.
result AlphaPruning prunes LLaMA-7B to 80% sparsity with reasonable perplexity.
Survey on reducing deep neural network training complexity.
problem Efficient training of large deep neural networks.
method Pruning and freezing parts of the network during training.
result Dimensionality reduction improves training efficiency.
Simplifies neural network compression with Gaussian priors and L1 regularization.
problem Neural network overfitting and scalability issues.
method Adds Gaussian priors and L1 regularization to the optimization problem for quantization and pruning.
result Achieves results competitive with state-of-the-art methods using simple modifications.
Dynamic pruning during training reduces deep network complexity without significant accuracy loss.
problem High memory and computational requirements of deep networks during training and inference.
method Dynamic pruning of convolutional filters during training, using L1 normalization for optimization.
result L1 normalization-based pruning yields up to 50% reduction in filters with minimal accuracy loss.
A new method prunes neural networks efficiently without losing effectiveness.
problem Efficient pruning of neural networks without sacrificing performance.
method Deterministic approximation of binary gates and L0 regularization. result Pruning neural networks significantly without loss in effectiveness.
EigenDamage reduces neural network size and FLOPs with structured pruning in the Kronecker-Factored Eigenbasis.
problem Reducing neural network size and FLOPs while maintaining accuracy for resource-constrained devices.
method Kronecker-Factored Eigenbasis reparameterization and Hessian-based structured pruning.
result Empirically validated improvements in model size and FLOPs with negligible accuracy loss.
FlipOut prunes neural networks by flipping weights' signs, achieving high sparsity.
problem Redundant weights in neural networks increase training time and resource usage.
method Uses sign flips during training to determine weight saliency for pruning.
result Competitive with existing methods, achieving state-of-the-art performance for high sparsity.
This paper analyzes privacy risks in neural network pruning and proposes a defense mechanism.
problem Privacy risks in neural network pruning due to membership inference attacks.
method Investigates the impact of pruning on prediction divergence and proposes a self-attention membership inference attack.
result Proposed defense mechanism mitigates privacy risks while maintaining sparsity and accuracy.