Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

18.8%37.5%56.3%75.0% · Jul 199319922001200920172026
48 results for embedding quantization

This work reduces model size by 86.11% for recommender systems using 4-bit quantization.

problem Large memory consumption in embedding vectors for recommender systems.
method Post-training 4-bit quantization on embedding tables, including row-wise uniform quantization and codebook-based quantization.
result Consistently reduces accuracy degradation while significantly reducing model size.

By the quantization condition compact quantizable Kaehler manifolds can be embedded into projective space. In this way they become projective varieties. The quantum Hilbert space of the Berezin-Toeplitz quantization (and of the geometric quantization) is the projective coordinate ring of the embedded manifold. This all…

2000-05-31abs ↗pdf ↗

By computing certain cohomology of Vect(M) of smooth vector fields we prove that on 1-dimensional manifolds M there is no quantization map intertwining the action of non-projective embeddings of the Lie algebra sl(2) into the Lie algebra Vect(M). Contrariwise, for projective embeddings sl(2)-equivariant quantization ex…

2006-01-14abs ↗pdf ↗

The abstract discusses embedding theorems for pseudo-Kähler manifolds.

problem Embedding theorems for pseudo-Kähler manifolds.
method Using quantizable pseudo-Kähler manifolds and Hermitian line bundles, the asymptotic expansion of Bergman kernels is analyzed.
result The asymptotic expansion of Bergman kernels implies analogues of Kodaira embedding theorem and Tian's almost-isometry theorem.

A new quantization strategy reduces Transformer model size and inference time.

problem Heavy computation load and memory overhead in Transformer models for mobile devices.
method Mixed precision quantization with varying bits per word in embedding blocks.
result 11.8x smaller model size and 3.5x speed up for on-device NMT.

This study optimizes quantized neural networks by considering model architecture and quantization types.

problem Optimizing quantized neural networks for low-power, high-throughput applications.
method Holistic approach including training methods and quantization-friendly architecture design.
result Deeper models are more sensitive to activation quantization, while wider models improve resilience to both weight and activation quantization.

A new algorithm for compressing latent representations in deep models.

problem Compressing continuous latent representations in deep models.
method Separates model design and training from quantization; uses adaptive quantization based on posterior uncertainty.
result Image compression with the proposed algorithm outperforms JPEG over a wide range of bit rates.

We investigate a quantization problem which asks for the construction of an algebra for relative elliptic problems of pseudodifferential type associated to smooth embeddings. Specifically, we study the problem for embeddings in the category of compact manifolds with corners. The construction of a calculus for elliptic …

2017-10-06abs ↗pdf ↗

This work optimizes quantization of linear models to reduce memory usage.

problem Optimizing memory usage for high-dimensional linear models.
method Information-theoretic framework with randomized embedding-based algorithms.
result Matching upper and lower bounds for minimax risk under quantization constraints.

Low-precision representation of deep neural networks (DNNs) is critical for efficient deployment of deep learning application on embedded platforms, however, converting the network to low precision degrades its performance. Crucially, networks that are designed for embedded applications usually suffer from increased de…

2019-06-07abs ↗pdf ↗

A fast binary embedding method preserves Euclidean distances in high-dimensional data.

problem Preserving Euclidean distances in high-dimensional datasets.
method Stable noise-shaping quantization of AxA x with AA a sparse Gaussian random matrix, followed by a linear transformation.
result Euclidean distances are approximated by the 1\ell_1 norm on binary sequences, leading to accurate binary codes.

In this lecture results on the Berezin-Toeplitz quantization of arbitrary compact quantizable Kaehler manifolds are presented. These results are obtained in joint work with M. Bordemann and E. Meinrenken. The existence of the Berezin-Toeplitz deformation quantization is also covered. Recent results obtained in joint wo…

2000-09-25abs ↗pdf ↗

Galen algorithm compresses neural networks for specific hardware with reduced latency.

problem Finding optimal compression policies for neural networks on specific hardware.
method Reinforcement learning using pruning and quantization to optimize inference latency.
result Compressed ResNet18 for ARM processor reduced inference latency by 80%.

Memory-augmented neural networks (MANNs) refer to a class of neural network models equipped with external memory (such as neural Turing machines and memory networks). These neural networks outperform conventional recurrent neural networks (RNNs) in terms of learning long-term dependency, allowing them to solve intrigui…

2017-11-10abs ↗pdf ↗

Efficient deep neural network (DNN) inference on mobile or embedded devices typically involves quantization of the network parameters and activations. In particular, mixed precision networks achieve better performance than networks with homogeneous bitwidth for the same size constraint. Since choosing the optimal bitwi…

2019-05-27abs ↗pdf ↗

Deep neural networks are the state-of-the-art methods for many real-world tasks, such as computer vision, natural language processing and speech recognition. For all its popularity, deep neural networks are also criticized for consuming a lot of memory and draining battery life of devices during training and inference.…

2018-08-13abs ↗pdf ↗

SGQuant reduces GNN memory usage without significant accuracy loss.

problem High memory consumption in GNNs limits their applicability on memory-constrained devices.
method Proposes a specialized GNN quantization scheme (SGQuant) with a quantization algorithm, fine-tuning scheme, and multi-granularity strategy.
result SGQuant reduces GNN memory footprint from 4.25x to 31.9x with minimal accuracy loss.

Currently, deep neural networks are deployed on low-power portable devices by first training a full-precision model using powerful hardware, and then deriving a corresponding low-precision model for efficient inference on such systems. However, training models directly with coarsely quantized weights is a key step towa…

2017-06-07abs ↗pdf ↗

Improved deep learning model deployment on tiny MCUs with mixed-precision quantization.

problem Memory limitations prevent accurate deployment of DNN models on tiny MCUs.
method Automated mixed-precision quantization using Reinforcement Learning for MCU constraints.
result Mixed-precision models achieve high accuracy with uniform quantization policies.

Just as semantic hashing can accelerate information retrieval, binary valued embeddings can significantly reduce latency in the retrieval of graphical data. We introduce a simple but effective model for learning such binary vectors for nodes in a graph. By imagining the embeddings as independent coin flips of varying b…

2018-03-25abs ↗pdf ↗

In this note we describe the recursion relations between two parameter HOMLFY and Kauffman polynomials of framed links These relation correspond to embeddings of quantized universal enveloping algebras. The relation corresponding to embeddings gngk×slnkg_{n}\supset g_{k}\times sl_{n-k} where gng_{n} is either so2n+1so_{2n+1}, $so…

2014-01-09abs ↗pdf ↗

New method accelerates CNNs for mobile devices by approximating tensors and quantizing weights.

problem Efficiently compress and accelerate CNNs for mobile devices.
method Low-rank tensor approximation in Tucker format combined with quantization of weights and activations.
result Our method significantly improves CNN performance on various classification tasks.

We consider existence and uniqueness of two kinds of coisotropic embeddings and deduce the existence of deformation quantizations of certain Poisson algebras of basic functions. First we show that any submanifold of a Poisson manifold satisfying a certain constant rank condition sits coisotropically inside some larger …

2006-11-15abs ↗pdf ↗

The ability to characterize the color content of natural imagery is an important application of image processing. The pixel by pixel coloring of images may be viewed naturally as points in color space, and the inherent structure and distribution of these points affords a quantization, through clustering, of the color i…

2012-02-20abs ↗pdf ↗

The paper defines and analyzes coherent and squeezed states on manifolds and their quantization.

problem Defining and characterizing coherent and squeezed states on various manifolds.
method Definition and analysis of Rawnsley-type coherent and squeezed states, Berezin quantization.
result Properties and quantization of coherent and squeezed states on manifolds.

Improved vector quantization using Gaussian mixtures for better codebook utilization.

problem Training instability and information loss in discrete vector quantization.
method Generalized vector quantization with Gaussian mixture model and aggregated categorical posterior evidence lower bound.
result GM-VQ improves codebook utilization and reduces information loss without heuristics.

G-Net constructs binary neural networks with high accuracy using randomized binary embeddings.

problem Creating high-accuracy binary neural networks with theoretical guarantees.
method Proposes a novel floating-point G-Net family with randomized binary embeddings and theoretical accuracy guarantees.
result Empirically, G-Net achieves almost 30% higher accuracy on CIFAR-10 compared to prior HDC models.

The paper explores maximal destabilizers for both K-stability and Chow-stability in unstable situations.

problem Exploring maximal destabilizers for K-stability and Chow-stability in unstable situations.
method Using non-Archimedean pluripotential theory and idealistic assumptions, the paper provides a route to show that maximal K-destabilizers are quantized by maximal Chow-destabilizers.
result Maximal K-destabilizers are quantized by maximal Chow-destabilizers.

In this thesis we study asymptotic behavior of projective embeddings of abelian varieties and their amoebas. The projective embeddings are given by theta functions. It is known that a Lagrangian fibration of the abelian variety determines a basis of theta functions. After reviewing the relation from the viewpoint of ge…

2006-04-14abs ↗pdf ↗