This work proposes a complete 8-bit quantization framework for large-scale deep neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper evaluates quantization techniques for deep learning inference.
We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations. Per-channel quantization of weights and per-layer quantization of activations to 8-bits of precision post-training produces classification accuracies within 2% of floating point networks…
AdaptivFloat improves deep learning inference accuracy at low precision.
Improved neural quantization reduces accuracy loss to less than 1% with 4-bit weights.
Improved deep learning model deployment on tiny MCUs with mixed-precision quantization.
This paper presents a novel end-to-end methodology for enabling the deployment of low-error deep networks on microcontrollers. To fit the memory and computational limitations of resource-constrained edge-devices, we exploit mixed low-bitwidth compression, featuring 8, 4 or 2-bit uniform quantization, and we model the i…
Improved machine translation with INT8 hardware using a novel training method.
FullyQT quantizes Transformer models to 8-bit precision without sacrificing translation quality.
Reduced precision computation for deep neural networks is one of the key areas addressing the widening compute gap driven by an exponential growth in model size. In recent years, deep learning training has largely migrated to 16-bit precision, with significant gains in performance and energy efficiency. However, attemp…
Low-precision DNNs have been extensively explored in order to reduce the size of DNN models for edge devices. Recently, the posit numerical format has shown promise for DNN data representation and compute with ultra-low precision in [5..8]-bits. However, previous studies were limited to studying posit for DNN inference…
The state-of-the-art hardware platforms for training Deep Neural Networks (DNNs) are moving from traditional single precision (32-bit) computations towards 16 bits of precision -- in large part due to the high energy efficiency and smaller bit storage associated with using reduced-precision representations. However, un…
VoiceFilter-Lite separates speech from background in real-time for on-device speech recognition.
Quantized Neural Networks (QNNs) are often used to improve network efficiency during the inference phase, i.e. after the network has been trained. Extensive research in the field suggests many different quantization schemes. Still, the number of bits required, as well as the best quantization scheme, are yet unknown. O…
WrapNet optimizes inference for low-resolution neural networks by using 8-bit additions.
We introduce a data-free quantization method for deep neural networks that does not require fine-tuning or hyperparameter selection. It achieves near-original model performance on common computer vision architectures and tasks. 8-bit fixed-point quantization is essential for efficient inference on modern deep learning …
Recently, the posit numerical format has shown promise for DNN data representation and compute with ultra-low precision ([5..8]-bit). However, majority of studies focus only on DNN inference. In this work, we propose DNN training using posits and compare with the floating point training. We evaluate on both MNIST and F…
In this paper, we firstly introduce a method to efficiently implement large-scale high-dimensional convolution with realistic memristor-based circuit components. An experiment verified simulator is adapted for accurate prediction of analog crossbar behavior. An improved conversion algorithm is developed to convert conv…
Low precision operations can provide scalability, memory savings, portability, and energy efficiency. This paper proposes SWALP, an approach to low precision training that averages low-precision SGD iterates with a modified learning rate schedule. SWALP is easy to implement and can match the performance of full-precisi…
Memory-augmented neural networks (MANNs) refer to a class of neural network models equipped with external memory (such as neural Turing machines and memory networks). These neural networks outperform conventional recurrent neural networks (RNNs) in terms of learning long-term dependency, allowing them to solve intrigui…
Implementing large-scale deep neural networks with high computational complexity on low-cost IoT devices may inevitably be constrained by limited computation resource, making the devices hard to respond in real-time. This disjunction makes the state-of-art deep learning algorithms, i.e. CNN (Convolutional Neural Networ…
Deep Neural Networks (DNNs) typically require massive amount of computation resource in inference tasks for computer vision applications. Quantization can significantly reduce DNN computation and storage by decreasing the bitwidth of network encodings. Recent research affirms that carefully selecting the quantization l…
GCWSNet improves neural network training speed and accuracy with power transformation.
Paper applies FloatSD8 to LSTM networks, reducing complexity and power.
New q-deformed integers help compute Jones polynomials efficiently.
Paper proposes a method to speed up DNNs by quantizing Winograd/Toom-Cook convolutions.
Generalized Steinberg module presentation for Gaussian and Eisenstein integers.
We present a new proof of Thurston's theorem that the unit ball of a seminorm on taking integer values on is a polyhedra defined by finitely many inequalities with integer coefficients.
Efficient Winograd convolution for INT8 networks using RNS.
Surgery obstructions extended to integer homology spheres using Heegaard Floer homology.
IDF++ improves integer discrete flows for lossless compression.
New links split by integer homology spheres but not by others.
We use Nathanson's -adic representation of integers to relate metric properties of Cayley graphs of the integers with respect to various infinite generating sets to problems in additive number theory. If consists of all powers of a fixed integer , we find explicit formulas for the smallest positive intege…
Geometric proof shows primes of form 3k+1 are norms of Eisenstein integers.
New method for probabilistic modeling of integer submodular functions.
An elementary proof shows that quasi-isometric groups to integers are virtually integers.
Study area-minimizing subgraphs in integer lattices.
New method finds lattice polygons that can be dissected into triangles with integer areas.
Recurrent Neural Networks (RNN) can be difficult to deploy on resource constrained devices due to their size.As a result, there is a need for compression techniques that can significantly compress RNNs without negatively impacting task accuracy. This paper introduces a method to compress RNNs for resource constrained e…
Neural networks with integer weights approximate continuous functions efficiently.
New examples show non-integer Hausdorff dimensions in collapsing spaces.
The integer hull of a polyhedron is the convex hull of the integer points contained in it. We show that the vertices of the integer hulls of a rational family of polyhedra of size O(n) have quasipolynomial coordinates. As a corollary, we show that the stable commutator length of elements in a surgery family is a ratio …
Paper uses integer programming for non-convex boosting in classification.
We propose a simple yet powerful framework for modeling integer-valued data, such as counts, scores, and rounded data. The data-generating process is defined by Simultaneously Transforming and Rounding (STAR) a continuous-valued process, which produces a flexible family of integer-valued distributions capable of modeli…
The article calculates a multiplying factor to convert rational Vassiliev invariants to integer-valued ones.
Improved RTM uses integer weights to reduce computation and increase interpretability.
Paper develops machine learning algorithms to learn optimal integer weights for clinical risk scores.
Groups of matrices with integer-like entries are studied.