Paper compresses RNNs for resource-constrained devices.
problem Difficulty deploying RNNs on resource-constrained devices.
method Uses Kronecker product (KP) to compress RNN layers.
result KP compresses RNN layers by 16-38x with minimal accuracy loss.
The paper defines and studies discrete p-density and compression-radius profiles of lattice knots.
problem Understanding geometric properties of lattice knots.
method Develops a framework for discrete p-density and compression-radius profiles of lattice knots, studying them on length-filtered sets and finite move-graph exploration.
result Density and compression-radius values are not monotone, illustrating distinct optimization problems.
This paper introduces matrix product state (MPS) decomposition as a new and systematic method to compress multidimensional data represented by higher-order tensors. It solves two major bottlenecks in tensor compression: computation and compression quality. Regardless of tensor order, MPS compresses tensors to matrices …
New method improves image compression using bits-back coding.
problem Lossy image compression with deep latent variable models.
method Iterative inference, stochastic annealing, bits-back coding.
result New state-of-the-art performance on lossy image compression.
SHVC improves image compression with fewer parameters.
problem Challenges in VAE compression, especially with bits-back coding.
method Introduces autoregressive sub-pixel convolution and autoregressive initial bits.
result Achieves state-of-the-art compression performance with fewer model parameters.
Model compresses event-like contexts using gated surprise signals.
problem Perceiving a dynamic world as organized events.
method Hierarchical, surprise-gated recurrent neural network architecture.
result Achieves best performance on multiple event processing tasks.
Develops a method for lossless compression using latent variable models.
problem Lossless compression of large datasets.
method Bits back with asymmetric numeral systems (BB-ANS) using latent variable models.
result Achieves state-of-the-art lossless compression of full-size colour images.
Single model corrects JPEG artifacts for various compression settings.
problem JPEG compression artifacts due to aggressive quantization.
method Parameterized architecture using quantization matrix.
result State-of-the-art performance across different quality settings.
In this paper, we consider linear state-space models with compressible innovations and convergent transition matrices in order to model spatiotemporally sparse transient events. We perform parameter and state estimation using a dynamic compressed sensing framework and develop an efficient solution consisting of two nes…
This study compresses BERT to make it suitable for low-capability devices.
problem Large Transformer models are resource-intensive.
method Survey and analysis of model compression techniques for BERT.
result Insights into best practices for compressing large-scale Transformer models.
Unified framework compresses GANs up to 47x with minimal quality loss.
problem High parameter complexity of GANs for resource-constrained devices.
method Unified optimization framework combining model distillation, channel pruning, and quantization.
result 47x compression of CartoonGAN with minimal quality degradation.
Generalizes bits back coding for time-series models with latent Markov structures.
problem Efficiently compressing time-series data with latent Markov structures.
method Extends bits back coding to time-series models with latent Markov structures, including HMMs and LGSSMs.
result Effective for small scale models, promising for larger scale settings like video compression.
Spectral methods reduce the complexity of Markov processes.
problem Modeling and simplifying state-transition systems.
method Spectral decomposition and state aggregation.
result Developed methods to estimate low-rank Markov models.
This paper improves the scalability of sparse neural network compression.
problem Sparse neural network compression for diverse data modalities.
method State-of-the-art sparsification techniques and meta-learning.
result Meta-learning sparse compression networks achieve new state-of-the-art results.
HiLLoC compresses large images losslessly using VAEs.
problem Lossless compression of large color photographs.
method Fully convolutional VAE models trained on ImageNet are applied to lossless compression.
result Achieves state-of-the-art compression for full-size ImageNet images.
This work tackles extractive compression by formulating it as tree transduction.
problem Extractive compression as a challenging natural language processing problem.
method Formulated as a parse tree transduction problem, using a deep neural model with Long Short-Term Memory extended to consider parent-child relationships.
result Achieves state-of-the-art performance on sentence compression benchmarks.
DiffC compresses images by diffusing Gaussian noise, outperforming state-of-the-art methods.
problem Efficient lossy image compression without transform coding.
method Unconditional diffusion generative models to encode and denoise corrupted pixels.
result DiffC outperforms HiFiC on ImageNet 64x64, achieving a 3 dB gain at high bitrates.
Adaptive Quantization Modules enable online continual compression of non-i.i.d data streams.
problem Learning to compress and store a dataset from a non-i.i.d data stream, only observing each sample once.
method Discrete auto-encoders and Adaptive Quantization Modules (AQM) to control compression ability.
result Significant gains on continual learning benchmarks with AQM replacing episodic memory.
TR-Nets compress deep networks by 11x for LeNet-5 and 243x for Wide ResNet.
problem Large neural networks require excessive memory and computation.
method Tensor Ring factorization to compress fully connected and convolutional layers.
result TR-Nets can compress LeNet-5 by 11x and Wide ResNet by 243x with minimal accuracy loss.
Optimizes MCMC output compression by selecting a subset of states.
problem Improving the efficiency of MCMC output compression.
method Greedy minimization of kernel Stein discrepancy for selecting a subset of states.
result The method provides an optimal subset of states for empirical approximation.
Compressive Transformer learns long-range sequences by compressing past memories.
problem Learning long-range sequences in language and speech models.
method Compressive Transformer compresses past memories for efficient long-range sequence learning.
result State-of-the-art performance on language and speech benchmarks.
NeLLoC improves image compression with parallel decoding.
problem Image compression with OOD generalization.
method Local autoregressive model with parallel decoding.
result Significant gains in compression runtime.
We extend quantization-aware training to extreme model compression.
problem Maximizing model accuracy with minimal model size.
method Quantize a random subset of weights during training, allowing unbiased gradients through other weights.
result Established new state-of-the-art compromises between accuracy and model size.
3LC compresses ML state changes for faster training, balancing accuracy and efficiency.
problem Efficiently compressing state changes in distributed ML to reduce communication time.
method Combines 3-value quantization, quartic encoding, and zero-run encoding to balance accuracy, compression, and computation overhead.
result Achieves up to 39-107X data compression ratio, maintains accuracy, and reduces training time by up to 23X.
Unified framework improves model compression while maintaining robustness.
problem Achieving high compression ratios without sacrificing adversarial robustness.
method Adversarially Trained Model Compression (ATMC) framework integrating multiple compression techniques.
result ATMC achieves better trade-off between model size, accuracy, and robustness.
Wavelets help compress neural networks efficiently.
problem Efficiently compressing linear layers in neural networks.
method Learnable wavelet transforms to compress RNNs.
result Wavelet compressed RNNs have fewer parameters and perform competitively.
New method improves neural network compression.
problem Efficiently compressing neural networks for mobile devices and inference.
method Combining Soft-Weight Sharing and Variational Dropout.
result New approach achieves state-of-the-art results in model compression.
Paper compresses RNNs using HT decomposition for better performance.
problem Large model sizes of RNNs in sequence analysis.
method Hierarchical Tucker (HT) tensor decomposition for model compression.
result HT-LSTM achieves better compression and accuracy than state-of-the-art methods.
DP-Net uses dynamic programming for efficient deep neural network compression.
problem Efficiently compressing deep neural networks while maintaining accuracy.
method Dynamic Programming for optimal weight quantization and clustering-friendly training.
result Achieves up to 77X compression ratio on Wide ResNet with minimal accuracy loss.
New method reduces neural image compression run-time by 50%.
problem Computational efficiency of neural image compression models.
method Automatic network optimization to reduce decoder complexity.
result Decreased decoder run-time by over 50%.
This paper compresses neural networks by permuting and quantizing weights.
problem Efficiently compressing large neural networks for resource-constrained platforms.
method Permuting and quantizing weights, connecting to rate-distortion theory, and using annealed quantization.
result Significant compression with minimal accuracy loss, e.g., 40-70% reduction in gap with uncompressed model.
Compressing neural nets is an active research problem, given the large size of state-of-the-art nets for tasks such as object recognition, and the computational limits imposed by mobile devices. We give a general formulation of model compression as constrained optimization. This includes many types of compression: quan…
Paper introduces compressibility loss for learning sparse neural network weights.
problem Learning highly compressible neural network weights.
method Applying a compressibility loss to minimize the negated sparsity of the signal.
result At critical points, weight vectors are ternary signals with a sparsity directly related to the objective value.
Improved image compression with diffusion models outperforming state-of-the-art methods.
problem Difficulties in replicating text-to-image success in image compression.
method Two-stage approach combining autoencoder targeting MSE followed by score-based decoder.
result Significantly improved perceptual quality at a given bit-rate, measured by FID score.
Paper introduces adversarial lossy compression for video artifacts reduction.
problem Unpleasant reconstruction artifacts in standard video coding schemes at low bit-rates.
method Adversarial lossy video compression model minimizing an adversarial distortion objective.
result Reduction of perceptual artifacts and detail reconstruction under extreme compression.
Iterative AutoML improves ASR model compression by 5x without WER degradation.
problem Challenges in achieving high compression levels without degrading ASR performance.
method Iterative AutoML-based Low Rank Factorization (LRF) approach.
result Achieved over 5x compression without WER degradation.
This paper compresses RNNs for IoT devices by 15-38x using Kronecker products.
problem Resource constraints on IoT devices make RNNs difficult to deploy.
method Kronecker product (KP) for compressing RNN layers by 15-38x with minimal accuracy loss.
result Kronecker product can compress RNNs by 50x when quantized to 8-bits.
Quantum TNCS uses machine learning to efficiently transmit data.
problem Efficient quantum communication of large datasets.
method Combining compressed sensing, tensor networks, and machine learning.
result High efficiency and accuracy in transmitting information.
Proposes a link between randomness and compression in deep learning.
problem Improving efficiency in deep learning training.
method Introduces a novel tomographic compression framework called Dual Tomographic Compression (DTC).
result Demonstrates high correlation between learning performance and Gibbs entropy over compression ratios.
New DNN boosts image retrieval efficiency.
problem Efficient compression of visual descriptors for large-scale image retrieval.
method Unsupervised multi-codebook quantization in a DNN architecture.
result Significantly outperforms existing methods on visual descriptor datasets.
A method extracts binary features directly from CS measurements for compressive image classification.
problem Efficiently classify images using compressive sensing without reconstruction.
method DCT-based approach for binary feature extraction from CS measurements, feature fusion with CNN features.
result Fused features outperform state-of-the-art methods in image classification.
HOTCAKE compresses CNNs by decomposing kernels into smaller parts.
problem Compressing deep CNNs without significant accuracy loss.
method Input channel decomposition, guided Tucker rank selection, higher order Tucker decomposition, fine-tuning.
result HOTCAKE produces highly compressed CNN models with good accuracy.
This paper proposes to perform authorship analysis using the Fast Compression Distance (FCD), a similarity measure based on compression with dictionaries directly extracted from the written texts. The FCD computes a similarity between two documents through an effective binary search on the intersection set between the …
A new method, REC, compresses images by encoding their latent representations efficiently.
problem Efficiently compressing single images with latent representations.
method Relative Entropy Coding (REC) that directly encodes latent representations with codelength close to relative entropy.
result REC is more efficient for single image compression compared to previous methods and is competitive for lossy compression.
DeepCMC compresses CSI for massive MIMO systems, reducing overhead and improving performance.
problem High CSI overhead in massive MIMO systems limits spectral efficiency.
method Deep learning-based fully convolutional neural network with residual layers and entropy coding.
result DeepCMC outperforms state-of-the-art schemes in CSI reconstruction quality for the same compression rate.
A simple method compresses neural network weights using entropy penalties.
problem Neural network weight compression for scalability and accuracy.
method Reparameterization with learned entropy penalty for compression.
result Maximized classification accuracy and model compressibility.
Paper models and compresses wideband CSI feedback in FDD MIMO systems.
problem Fundamental limits of channel state information (CSI) feedback in FDD massive MIMO systems.
method Modeling CSI as a Gaussian-mixture source with latent geometry states, proposing Gaussian-mixture transform coding (GMTC).
result Near-optimal CSI compression achieved through state-adaptive transform coding without large neural encoders.
New video compression method outperforms traditional approaches.
problem Efficient video compression in low latency mode.
method Feedback Recurrent Autoencoder network architecture.
result State of the art MS-SSIM/rate performance on UVG dataset.