Compressing word embeddings is important for deploying NLP models in memory-constrained settings. However, understanding what makes compressed embeddings perform well on downstream tasks is challenging---existing measures of compression quality often fail to distinguish between embeddings that perform well and those th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes IIQ for compressing embedding vectors.
The recent advances in deep neural networks (DNNs) make them attractive for embedded systems. However, it can take a long time for DNNs to make an inference on resource-constrained computing devices. Model compression techniques can address the computation issue of deep inference on embedded devices. This technique is …
G-CREWE efficiently aligns large networks using node embeddings and compression.
Compressed LLM embeddings improve noisy regression tasks without overfitting.
Given Poincare spaces M and X, we study the possibility of compressing embeddings of M x I in X x I down to embeddings of M in X. This results in a new approach to embedding in the metastable range both in the smooth and Poincare duality categories.
This is the second of three papers about the Compression Theorem. We give proofs of Gromov's theorem on directed embeddings [M Gromov, Partial differential relations, Springer--Verlag (1986); 2.4.5 C'] and of the Normal Deformation Theorem [The compression theorem I; 4.7], arxiv:math.GT/9712235.
Improves compression of neural networks for embedded systems.
Production recommendation systems rely on embedding methods to represent various features. An impeding challenge in practice is that the large embedding matrix incurs substantial memory footprint in serving as the number of features grows over time. We propose a similarity-aware embedding matrix compression method call…
Deep learning models have become state of the art for natural language processing (NLP) tasks, however deploying these models in production system poses significant memory constraints. Existing compression methods are either lossy or introduce significant latency. We propose a compression method that leverages low rank…
Embedding layers are commonly used to map discrete symbols into continuous embedding vectors that reflect their semantic meanings. Despite their effectiveness, the number of parameters in an embedding layer increases linearly with the number of symbols and poses a critical challenge on memory and storage constraints. I…
If one tries to embed a metric space uniformly in Hilbert space, how close to quasi-isometric could the embedding be? We answer this question for finite dimensional CAT(0) cube complexes and for hyperbolic groups. In particular, we show that the Hilbert space compression of any hyperbolic group is 1.
We develop embeddings for nonlinear subspaces preserving vector norms.
Galen algorithm compresses neural networks for specific hardware with reduced latency.
We produce embeddings of knots in thin position that admit compressible thin levels. We also find the bridge number of tangle sums where each tangle is high distance.
Model compression is essential for serving large deep neural nets on devices with limited resources or applications that require real-time responses. As a case study, a state-of-the-art neural language model usually consists of one or more recurrent layers sandwiched between an embedding layer used for representing inp…
Deep neural networks (DNNs) have been quite successful in solving many complex learning problems. However, DNNs tend to have a large number of learning parameters, leading to a large memory and computation requirement. In this paper, we propose a model compression framework for efficient training and inference of deep …
New approach for sharing deep learning costs between devices and cloud.
A new method compresses conditional distributions of labelled data.
The topological information is essential for studying the relationship between nodes in a network. Recently, Network Representation Learning (NRL), which projects a network into a low-dimensional vector space, has been shown their advantages in analyzing large-scale networks. However, most existing NRL methods are desi…
New method alters Seifert surfaces without compression.
A new algorithm for compressing latent representations in deep models.
Study 1-bit compressive sensing with generative models, improving recovery accuracy.
Diffusion models improve image compression at low bit-rates.
This dissertation improves classical compression techniques using deep learning.
Although the convolutional neural networks (CNNs) have become popular for various image processing and computer vision task recently, it remains a challenging problem to reduce the storage cost of the parameters for resource-limited platforms. In the previous studies, tensor decomposition (TD) has achieved promising co…
We show that the type function of a space with finite asymptotic dimension estimates its Hilbert (or any ) compression. The method allows to obtain the lower bound of the compression of the lamplighter group , which has infinite asymptotic dimension.
New method compresses neural networks up to 14x with minimal performance loss.
ALF reduces network parameters and operations by 70% and 61%, respectively, on embedded hardware.
One of the earliest conjectures in computational learning theory-the Sample Compression conjecture-asserts that concept classes (equivalently set systems) admit compression schemes of size linear in their VC dimension. To-date this statement is known to be true for maximum classes---those that possess maximum cardinali…
Bayesian sparsification improves complex-valued neural networks by 50-100x with minimal performance loss.
A new method compresses NLP networks by using multiple subspaces instead of a single one.
We give first examples of finitely generated groups having an intermediate, with values in (0,1), Hilbert space compression (which is a numerical parameter measuring the distortion required to embed a metric space into Hilbert space). These groups include certain diagram groups. In particular, we show that the Hilbert …
Paper proposes Nyström sketches for better adaptive compressive learning.
Diffusion maps are a commonly used kernel-based method for manifold learning, which can reveal intrinsic structures in data and embed them in low dimensions. However, as with most kernel methods, its implementation requires a heavy computational load, reaching up to cubic complexity in the number of data points. This l…
A new embedding method for high-dimensional data.
This paper deals with two related problems, namely distance-preserving binary embeddings and quantization for compressed sensing . First, we propose fast methods to replace points from a subset , associated with the Euclidean metric, with points in the cube and we associa…
This the first of a set of three papers about the Compression Theorem: if M^m is embedded in Q^q X R with a normal vector field and if q-m > 0, then the given vector field can be straightened (ie, made parallel to the given R direction) by an isotopy of M and normal field in Q X R. The theorem can be deduced from Gromo…
Motivated by the ever-increasing demands for limited communication bandwidth and low-power consumption, we propose a new methodology, named joint Variational Autoencoders with Bernoulli mixture models (VAB), for performing clustering in the compressed data domain. The idea is to reduce the data dimension by Variational…
With the development of deep neural networks, the size of network models becomes larger and larger. Model compression has become an urgent need for deploying these network models to mobile or embedded devices. Model quantization is a representative model compression technique. Although a lot of quantization methods hav…
We propose and evaluate new techniques for compressing and speeding up dense matrix multiplications as found in the fully connected and recurrent layers of neural networks for embedded large vocabulary continuous speech recognition (LVCSR). For compression, we introduce and study a trace norm regularization technique f…
New tool: relative Hopf invariant for Poincaré surgery.
Multi-criteria recommender systems have been increasingly valuable for helping consumers identify the most relevant items based on different dimensions of user experiences. However, previously proposed multi-criteria models did not take into account latent embeddings generated from user reviews, which capture latent se…
Deep neural networks (DNNs) have been expanded into medical fields and triggered the revolution of some medical applications by extracting complex features and achieving high accuracy and performance, etc. On the contrast, the large-scale network brings high requirements of both memory storage and computation resource,…
In natural language processing, a lot of the tasks are successfully solved with recurrent neural networks, but such models have a huge number of parameters. The majority of these parameters are often concentrated in the embedding layer, which size grows proportionally to the vocabulary length. We propose a Bayesian spa…
A new method uses a frozen language model to improve sample efficiency in reinforcement learning.
IDF++ improves integer discrete flows for lossless compression.
New faster, space-saving methods for subspace embeddings in tensors.