Low-bit training framework reduces energy consumption in CNNs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper develops a statistical framework for quantized training of deep neural networks.
Higher-order tensors can represent scores in a rating system, frames in a video, and images of the same subject. In practice, the measurements are often highly quantized due to the sampling strategies or the quality of devices. Existing works on tensor recovery have focused on data losses and random noises. Only a few …
New method uses tensor networks to price multi-asset options efficiently.
Quantization and reduction studied for CR manifolds with group actions.
This paper finds a new way to compress CNN weights, improving on pruning and quantization.
We prove the existence and the uniqueness of a conformally equivariant symbol calculus and quantization on any conformally flat pseudo-Riemannian manifold $(M,\rg)$. In other words, we establish a canonical isomorphism between the spaces of polynomials on and of differential operators on tensor densities over $M…
Kernel Quantization improves CNN compression without sacrificing performance.
New method accelerates CNNs for mobile devices by approximating tensors and quantizing weights.
Verma Howe duality connects tensor products of Verma modules to LKB representations.
The article defines and compares two types of quantizations on compact manifolds.
Quantization-aware training can recover accuracy lost by post-training quantization.
We investigate the concept of projectively equivariant quantization in the framework of super projective geometry. When the projective superalgebra pgl(p+1|q) is simple, our result is similar to the classical one in the purely even case: we prove the existence and uniqueness of the quantization except in some critical …
We consider natural differential operations acting on sections of tensor vector bundles. Arrising problems can be reformulated as invariant theoretical problems (the IT-reduction). We give examples of usage of the IT-reduction. In particular, on a manifold with a connection and a Poisson structure we construct the cano…
BatchNorm helps train quantized networks by avoiding gradient explosion.
Optimal gradient quantization reduces communication costs in distributed deep learning.
StatQAT optimizes quantization for deep networks, reducing computational cost and memory usage.
Proposes a robust neural network quantization method.
We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations. Per-channel quantization of weights and per-layer quantization of activations to 8-bits of precision post-training produces classification accuracies within 2% of floating point networks…
New method trains quantized neural networks to global optimality.
We analyze the effect of quantizing weights and activations of neural networks on their loss and derive a simple regularization scheme that improves robustness against post-training quantization. By training quantization-ready networks, our approach enables storing a single set of weights that can be quantized on-deman…
In this paper we give a construction of Fedosov quantization incorporating the odd variables and an analogous formula to Getzler's pseudodifferential calculus composition formula is obtained. A Fedosov type connection is constructed on the bundle of Weyl tensor Clifford algebras over the cotangent bundle of a Riemannia…
Quantized Adam reduces communication cost in deep learning training.
For a symplectic manifold with quantizing line bundle, a choice of almost complex structure determines a Laplacian acting on tensor powers of the bundle. For high tensor powers Guillemin-Uribe showed that there is a well-defined cluster of low-lying eigenvalues, whose distribution is described by a spectral density fun…
This paper introduces a differentiable, scalable quantization method for neural networks.
Neural network quantization procedure is the necessary step for porting of neural networks to mobile devices. Quantization allows accelerating the inference, reducing memory consumption and model size. It can be performed without fine-tuning using calibration procedure (calculation of parameters necessary for quantizat…
Let be an arbitrary complex manifold and let be a Hermitian holomorphic line bundle over . We introduce the Berezin-Toeplitz quantization of the open set of where the curvature on is non-degenerate. The quantum spaces are the spectral spaces corresponding to ( fixed), of the Kodaira…
Post-training quantization method using multiple low-precision points achieves higher precision for critical weights.
Geometric quantization extended to big line bundles.
Bayesian Bits unifies quantization and pruning through gradient optimization.
In this paper, we consider several compression techniques for the language modeling problem based on recurrent neural networks (RNNs). It is known that conventional RNNs, e.g, LSTM-based networks in language modeling, are characterized with either high space complexity or substantial inference time. This problem is esp…
Currently, deep neural networks are deployed on low-power portable devices by first training a full-precision model using powerful hardware, and then deriving a corresponding low-precision model for efficient inference on such systems. However, training models directly with coarsely quantized weights is a key step towa…
Unified finetuning of all quantization degrees of freedom achieves state-of-the-art 4-bit quantization.
The deep layers of modern neural networks extract a rather rich set of features as an input propagates through the network. This paper sets out to harvest these rich intermediate representations for quantization with minimal accuracy loss while significantly reducing the memory footprint and compute intensity of the DN…
For a very ample line bundle L on a compact connected complex manifold X, with a real structure, we discuss entanglement properties of certain sequences of vectors in tensor products of spaces of holomorphic sections of powers of L.
Deep quantization of neural networks (below eight bits) offers significant promise in reducing their compute and storage cost. Albeit alluring, without special techniques for training and optimization, deep quantization results in significant accuracy loss. To further mitigate this loss, we propose a novel sinusoidal r…
Improved deep learning model deployment on tiny MCUs with mixed-precision quantization.
The paper constructs quantizations for symplectic manifolds with specific Laplacian properties.
Deep neural network (DNN) quantization converting floating-point (FP) data in the network to integers (INT) is an effective way to shrink the model size for memory saving and simplify the operations for compute acceleration. Recently, researches on DNN quantization develop from inference to training, laying a foundatio…
Tensor factorization has become an increasingly popular approach to knowledge graph completion(KGC), which is the task of automatically predicting missing facts in a knowledge graph. However, even with a simple model like CANDECOMP/PARAFAC(CP) tensor decomposition, KGC on existing knowledge graphs is impractical in res…
Proposes QEP to mitigate quantization error propagation in layer-wise post-training quantization.
Neural network models are resource hungry. It is difficult to deploy such deep networks on devices with limited resources, like smart wearables, cellphones, drones, and autonomous vehicles. Low bit quantization such as binary and ternary quantization is a common approach to alleviate this resource requirements. Ternary…
We extend quantization-aware training to extreme model compression.
We generalize the results of Montgomery for the Bochner Laplacian on high tensor powers of a line bundle. When specialized to Riemann surfaces, this leads to the Bergman kernel expansion and geometric quantization results for semi-positive line bundles whose curvature vanishes at finite order. The proof exploits the re…
Low bit-width integer weights and activations are very important for efficient inference, especially with respect to lower power consumption. We propose Monte Carlo methods to quantize the weights and activations of pre-trained neural networks without any re-training. By performing importance sampling we obtain quantiz…
This article is a survey of recent work of the authors developing a new approach to quantization based on the equivariance with respect to some Lie group of symmetries. Examples are provided by conformal and projective differential geometry: given a smooth manifold M endowed with a flat conformal/projective structure, …
AdaRound improves post-training quantization of neural networks.
We investigate the concept of equivariant quantization over the superspace R^{p+q|2r}, with respect to the orthosymplectic algebra osp(p+1,q+1|2r). Our methods and results vary upon the superdimension p+q-2r. When the superdimension is nonzero, we manage to obtain a result which is similar to the classical theorem of D…