NESTA accelerates neural networks by compressing Hamming weights.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We revisit fuzzy neural network with a cornerstone notion of generalized hamming distance, which provides a novel and theoretically justified framework to re-interpret many useful neural network techniques in terms of fuzzy logic. In particular, we conjecture and empirically illustrate that, the celebrated batch normal…
In this paper, we present a novel approach for fine-tuning a decoder-side neural network in the context of image compression, such that the weight-updates are better compressible. At encoder side, we fine-tune a pre-trained artifact removal network on target data by using a compression objective applied on the weight-u…
Note proves bi-Lipschitz embeddings for Ham(S^2) metrics.
Model compression has been introduced to reduce the required hardware resources while maintaining the model accuracy. Lots of techniques for model compression, such as pruning, quantization, and low-rank approximation, have been suggested along with different inference implementation characteristics. Adopting model com…
In this paper we apply a compressibility loss that enables learning highly compressible neural network weights. The loss was previously proposed as a measure of negated sparsity of a signal, yet in this paper we show that minimizing this loss also enforces the non-zero parts of the signal to have very low entropy, thus…
Kernel Quantization improves CNN compression without sacrificing performance.
This work compresses heavy-tailed weight matrices for tighter generalization bounds.
With the development of deep neural networks, the size of network models becomes larger and larger. Model compression has become an urgent need for deploying these network models to mobile or embedded devices. Model quantization is a representative model compression technique. Although a lot of quantization methods hav…
We prove that contains an infinite cyclic subgroup, where is the Hamiltonian group of the one point blow up of . We give a sufficient condition for the group to contain an infinite cyclic subgroup, when is a general toric manifold.
We verify here some variants of topological and dynamical flavor of the injectivity radius conjecture in Hofer geometry, Lalonde-Savelyev \cite{citeLalondeSavelyevOntheinjectivityradiusinHofergeometry} in the case of and , for a closed positive genus surface. In particular we show that any lo…
Weight Squeezing transfers knowledge from large models to smaller ones, improving performance and speed.
The success of deep learning in numerous application domains created the de- sire to run and train them on mobile devices. This however, conflicts with their computationally, memory and energy intense nature, leading to a growing interest in compression. Recent work by Han et al. (2015a) propose a pipeline that involve…
Due to a resource-constrained environment, network compression has become an important part of deep neural networks research. In this paper, we propose a new compression method, \textit{Inter-Layer Weight Prediction} (ILWP) and quantization method which quantize the predicted residuals between the weights in all convol…
We propose using five data-driven community detection approaches from social networks to partition the label space for the task of multi-label classification as an alternative to random partitioning into equal subsets as performed by RAkELd: modularity-maximizing fastgreedy and leading eigenvector, infomap, walktrap an…
Paper improves MIRACLE for faster, more robust neural network compression.
We compress large neural networks for quick adaptation to specific contexts.
The large memory requirements of deep neural networks limit their deployment and adoption on many devices. Model compression methods effectively reduce the memory requirements of these models, usually through applying transformations such as weight pruning or quantization. In this paper, we present a novel scheme for l…
Model compression has gained a lot of attention due to its ability to reduce hardware resource requirements significantly while maintaining accuracy of DNNs. Model compression is especially useful for memory-intensive recurrent neural networks because smaller memory footprint is crucial not only for reducing storage re…
This paper investigates compression techniques for deep neural networks to reduce their size without sacrificing performance.
Let be a compact oriented surface. We construct homogeneous quasimorphisms on , on and on generalizing the constructions of Gambaudo-Ghys and Polterovich. We prove that there are infinitely many linearly independent homogeneous quasimorphisms on , on $Diff_0(…
DP-Net uses dynamic programming for efficient deep neural network compression.
We consider a decomposition method for compressive streaming data in the context of online compressive Robust Principle Component Analysis (RPCA). The proposed decomposition solves an - cluster-weighted minimization to decompose a sequence of frames (or vectors), into sparse and low-rank components, from com…
Develops homotopies for Lagrangian field theory using advanced algebraic structures.
We present a lower bound for a fragmentation norm and construct a bi-Lipschitz embedding with respect to the fragmentation norm on the group of Hamiltonian diffeomorphisms of a symplectic manifold . As an application, we provide an answer to Brandenbursk…
New compression methods handle biased input sequences for more accurate posterior summaries.
Differentiable model compression adds noise to parameters during training.
This paper studies how to compress neural networks while maintaining accuracy.
We extend quantization-aware training to extreme model compression.
Vectors of data are at the heart of machine learning and data mining. Recently, vector quantization methods have shown great promise in reducing both the time and space costs of operating on vectors. We introduce a vector quantization algorithm that can compress vectors over 12x faster than existing techniques while al…
We consider the problem of deep neural net compression by quantization: given a large, reference net, we want to quantize its real-valued weights using a codebook with entries so that the training loss of the quantized net is minimal. The codebook can be optimally learned jointly with the net, or fixed, as for bina…
Bayesian neural networks are compressed using feature and weight pruning based on posterior inclusion probabilities.
DACE estimates covariance from compressed data, improving accuracy.
Proposes a link between randomness and compression in deep learning.
This paper proposes a new method for efficient data compression using Bayesian neural networks.
The current trend of pushing CNNs deeper with convolutions has created a pressing demand to achieve higher compression gains on CNNs where convolutions dominate the computation and parameter amount (e.g., GoogLeNet, ResNet and Wide ResNet). Further, the high energy consumption of convolutions limits its deployment on m…
We describe a simple and general neural network weight compression approach, in which the network parameters (weights and biases) are represented in a "latent" space, amounting to a reparameterization. This space is equipped with a learned probability model, which is used to impose an entropy penalty on the parameter r…
Theoretical framework for neural network compression using sparsity norms.
Paper presents a new trie for integer sketches to improve similarity searches.
Model compression techniques, such as pruning and quantization, are becoming increasingly important to reduce the memory footprints and the amount of computations. Despite model size reduction, achieving performance enhancement on devices is, however, still challenging mainly due to the irregular representations of spa…
We apply Gromov's ham sandwich method to get (1) domain monotonicity (up to a multiplicative constant factor); (2) reverse domain monotonicity (up to a multiplicative constant factor); and (3) universal inequalities for Neumann eigenvalues of the Laplacian on bounded convex domains in a Euclidean space.
This paper finds a new way to compress CNN weights, improving on pruning and quantization.
We present a novel method of compression of deep Convolutional Neural Networks (CNNs) by weight sharing through a new representation of convolutional filters. The proposed method reduces the number of parameters of each convolutional layer by learning a 1D vector termed Filter Summary (FS). The convolutional filters ar…
IMPACT optimizes LLM compression by focusing on activation importance, reducing model size up to 55.4%.
New STH distance finds patterns in event timeseries without resampling.
Layer fusion reduces deep neural network layers with minimal loss in accuracy.
Three new efficient algorithms project vectors onto weighted l1 ball.
This paper deals with two related problems, namely distance-preserving binary embeddings and quantization for compressed sensing . First, we propose fast methods to replace points from a subset , associated with the Euclidean metric, with points in the cube and we associa…