This work reduces model size by 86.11% for recommender systems using 4-bit quantization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes IIQ for compressing embedding vectors.
By the quantization condition compact quantizable Kaehler manifolds can be embedded into projective space. In this way they become projective varieties. The quantum Hilbert space of the Berezin-Toeplitz quantization (and of the geometric quantization) is the projective coordinate ring of the embedded manifold. This all…
By computing certain cohomology of Vect(M) of smooth vector fields we prove that on 1-dimensional manifolds M there is no quantization map intertwining the action of non-projective embeddings of the Lie algebra sl(2) into the Lie algebra Vect(M). Contrariwise, for projective embeddings sl(2)-equivariant quantization ex…
The article defines and compares two types of quantizations on compact manifolds.
Quantized neural networks are vulnerable to adversarial attacks.
Study on quantum particle evolution on Grushin cylinder, embedding in R^3.
Computational aspects of the optimal consumption and investment with the partially observed stochastic volatility of the asset prices are considered. The new quantization approach to filtering - density quantization - is introduced which reduces the original infinite dimensional state space of the problem to the finite…
This paper deals with two related problems, namely distance-preserving binary embeddings and quantization for compressed sensing . First, we propose fast methods to replace points from a subset , associated with the Euclidean metric, with points in the cube and we associa…
The abstract discusses embedding theorems for pseudo-Kähler manifolds.
A new quantization strategy reduces Transformer model size and inference time.
This study optimizes quantized neural networks by considering model architecture and quantization types.
A new algorithm for compressing latent representations in deep models.
We investigate a quantization problem which asks for the construction of an algebra for relative elliptic problems of pseudodifferential type associated to smooth embeddings. Specifically, we study the problem for embeddings in the category of compact manifolds with corners. The construction of a calculus for elliptic …
With the development of deep neural networks, the size of network models becomes larger and larger. Model compression has become an urgent need for deploying these network models to mobile or embedded devices. Model quantization is a representative model compression technique. Although a lot of quantization methods hav…
Embedding layers are commonly used to map discrete symbols into continuous embedding vectors that reflect their semantic meanings. Despite their effectiveness, the number of parameters in an embedding layer increases linearly with the number of symbols and poses a critical challenge on memory and storage constraints. I…
This work optimizes quantization of linear models to reduce memory usage.
Low-precision representation of deep neural networks (DNNs) is critical for efficient deployment of deep learning application on embedded platforms, however, converting the network to low precision degrades its performance. Crucially, networks that are designed for embedded applications usually suffer from increased de…
PoET-BiN reduces power consumption in neural networks on embedded devices.
Cyber-Physical Systems (CPSs) have been pervasive including smart grid, autonomous automobile systems, medical monitoring, process control systems, robotics systems, and automatic pilot avionics. As usually implemented on embedded devices, CPS is typically constrained by computation capacity and energy consumption. In …
A fast binary embedding method preserves Euclidean distances in high-dimensional data.
In this lecture results on the Berezin-Toeplitz quantization of arbitrary compact quantizable Kaehler manifolds are presented. These results are obtained in joint work with M. Bordemann and E. Meinrenken. The existence of the Berezin-Toeplitz deformation quantization is also covered. Recent results obtained in joint wo…
In the first part of the paper we describe the complex geometry of the universal Teichmüller space , which may be realized as an open subset in the complex Banach space of holomorphic quadratic differentials in the unit disc. The quotient of the diffeomorphism group of the circle modulo Möbius …
M5-branes' flux quantization linked to non-abelian cohomology.
Galen algorithm compresses neural networks for specific hardware with reduced latency.
Memory-augmented neural networks (MANNs) refer to a class of neural network models equipped with external memory (such as neural Turing machines and memory networks). These neural networks outperform conventional recurrent neural networks (RNNs) in terms of learning long-term dependency, allowing them to solve intrigui…
Efficient deep neural network (DNN) inference on mobile or embedded devices typically involves quantization of the network parameters and activations. In particular, mixed precision networks achieve better performance than networks with homogeneous bitwidth for the same size constraint. Since choosing the optimal bitwi…
Deep neural networks are the state-of-the-art methods for many real-world tasks, such as computer vision, natural language processing and speech recognition. For all its popularity, deep neural networks are also criticized for consuming a lot of memory and draining battery life of devices during training and inference.…
SGQuant reduces GNN memory usage without significant accuracy loss.
Currently, deep neural networks are deployed on low-power portable devices by first training a full-precision model using powerful hardware, and then deriving a corresponding low-precision model for efficient inference on such systems. However, training models directly with coarsely quantized weights is a key step towa…
Improved deep learning model deployment on tiny MCUs with mixed-precision quantization.
Just as semantic hashing can accelerate information retrieval, binary valued embeddings can significantly reduce latency in the retrieval of graphical data. We introduce a simple but effective model for learning such binary vectors for nodes in a graph. By imagining the embeddings as independent coin flips of varying b…
Deep learning as a means to inferencing has proliferated thanks to its versatility and ability to approach or exceed human-level accuracy. These computational models have seemingly insatiable appetites for computational resources not only while training, but also when deployed at scales ranging from data centers all th…
In this note we describe the recursion relations between two parameter HOMLFY and Kauffman polynomials of framed links These relation correspond to embeddings of quantized universal enveloping algebras. The relation corresponding to embeddings where is either , $so…
New method accelerates CNNs for mobile devices by approximating tensors and quantizing weights.
We consider existence and uniqueness of two kinds of coisotropic embeddings and deduce the existence of deformation quantizations of certain Poisson algebras of basic functions. First we show that any submanifold of a Poisson manifold satisfying a certain constant rank condition sits coisotropically inside some larger …
Efficient neural networks for resource-constrained systems.
Quantization of neural networks has become common practice, driven by the need for efficient implementations of deep neural networks on embedded devices. In this paper, we exploit an oft-overlooked degree of freedom in most networks - for a given layer, individual output channels can be scaled by any factor provided th…
The ability to characterize the color content of natural imagery is an important application of image processing. The pixel by pixel coloring of images may be viewed naturally as points in color space, and the inherent structure and distribution of these points affords a quantization, through clustering, of the color i…
The paper defines and analyzes coherent and squeezed states on manifolds and their quantization.
Improved vector quantization using Gaussian mixtures for better codebook utilization.
Deep learning models have become state of the art for natural language processing (NLP) tasks, however deploying these models in production system poses significant memory constraints. Existing compression methods are either lossy or introduce significant latency. We propose a compression method that leverages low rank…
Separable Bregman divergences induce Riemannian metric spaces that are isometric to the Euclidean space after monotone embeddings. We investigate fixed rate quantization and its codebook Voronoi diagrams, and report on experimental performances of partition-based, hierarchical, and soft clustering algorithms with respe…
G-Net constructs binary neural networks with high accuracy using randomized binary embeddings.
For particles constrained on a curved surface, how to perform quantization within Dirac's canonical quantization scheme is a long-standing problem. On one hand, Dirac stressed that the Cartesian coordinate system has fundamental importance in passing from the classical Hamiltonian to its quantum mechanical form while p…
The paper explores maximal destabilizers for both K-stability and Chow-stability in unstable situations.
Quantum traces embed into quantum tori for surface skein algebras.
In this thesis we study asymptotic behavior of projective embeddings of abelian varieties and their amoebas. The projective embeddings are given by theta functions. It is known that a Lagrangian fibration of the abelian variety determines a basis of theta functions. After reviewing the relation from the viewpoint of ge…