NeLLoC improves image compression with parallel decoding.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Neural networks compress uninformative input directions, improving test error.
SHVC improves image compression with fewer parameters.
Paper proposes DCT for efficient hybrid parallel training of large recommendation models.
Clapping reduces memory usage in distributed optimization by reusing data samples.
ARDMs are a new model class for autoregressive diffusion that generalize existing models and can compress data efficiently.
Adaptive quantization improves SGD accuracy in data-parallel settings.
PALMS reconstructs large-scale networks efficiently with parallel computing.
This study presents a new lossy image compression method that utilizes the multi-scale features of natural images. Our model consists of two networks: multi-scale lossy autoencoder and parallel multi-scale lossless coder. The multi-scale lossy autoencoder extracts the multi-scale image features to quantized variables a…
Pruning is an efficient model compression technique to remove redundancy in the connectivity of deep neural networks (DNNs). Computations using sparse matrices obtained by pruning parameters, however, exhibit vastly different parallelism depending on the index representation scheme. As a result, fine-grained pruning ha…
A new gradient quantization scheme improves communication efficiency in distributed training.
As the size and complexity of models and datasets grow, so does the need for communication-efficient variants of stochastic gradient descent that can be deployed to perform parallel model training. One popular communication-compression method for data-parallel SGD is QSGD (Alistarh et al., 2017), which quantizes and en…
Data parallelism can boost the training speed of convolutional neural networks (CNN), but could suffer from significant communication costs caused by gradient aggregation. To alleviate this problem, several scalar quantization techniques have been developed to compress the gradients. But these techniques could perform …
Tensor decomposition on big data has attracted significant attention recently. Among the most popular methods is a class of algorithms that leverages compression in order to reduce the size of the tensor and potentially parallelize computations. A fundamental requirement for such methods to work properly is that the lo…
APMSqueeze improves Adam for faster training with less communication.
We study gradient compression methods to alleviate the communication bottleneck in data-parallel distributed optimization. Despite the significant attention received, current compression schemes either do not scale well or fail to achieve the target test accuracy. We propose a new low-rank gradient compressor based on …
We prove a conjecture of Menasco and Zhang that if a tangle is completely tubing compressible then it consists of at most two families of parallel strands. This is related to problems of graphs in 3-manifold. A 1-vertex graph in a 3-manifold with a genus 1 Heegaard splitting is standard if it consists of one or…
Paper introduces a new adaptive gradient method with gradient compression for distributed training.
Deep latent variable models have seen recent success in many data domains. Lossless compression is an application of these models which, despite having the potential to be highly useful, has yet to be implemented in a practical manner. We present `Bits Back with ANS' (BB-ANS), a scheme to perform lossless compression w…
Discrete Flow Maps bypass sequential prediction limits for parallel text generation.
Deep Gaussian processes provide a flexible approach to probabilistic modelling of data using either supervised or unsupervised learning. For tractable inference approximations to the marginal likelihood of the model must be made. The original approach to approximate inference in these models used variational compressio…
Paper proposes ADC framework to reduce ViT SL training communication overhead.
New methods reduce communication in distributed training for variational inequalities.
Modern deep learning models are often trained in parallel over a collection of distributed machines to reduce training time. In such settings, communication of model updates among machines becomes a significant performance bottleneck and various lossy update compression techniques have been proposed to alleviate this p…
Multiresolution Matrix Factorization (MMF) was recently introduced as a method for finding multiscale structure and defining wavelets on graphs/matrices. In this paper we derive pMMF, a parallel algorithm for computing the MMF factorization. Empirically, the running time of pMMF scales linearly in the dimension for spa…
Highly distributed training of Deep Neural Networks (DNNs) on future compute platforms (offering 100 of TeraOps/s of computational capacity) is expected to be severely communication constrained. To overcome this limitation, new gradient compression techniques are needed that are computationally friendly, applicable to …
Deep neural networks (DNNs) have been quite successful in solving many complex learning problems. However, DNNs tend to have a large number of learning parameters, leading to a large memory and computation requirement. In this paper, we propose a model compression framework for efficient training and inference of deep …
Nonparametric regression for massive numbers of samples (n) and features (p) is an increasingly important problem. In big n settings, a common strategy is to partition the feature space, and then separately apply simple models to each partition set. We propose an alternative approach, which avoids such partitioning and…
Improved Bayesian optimization for conditional parameter spaces.
DORE reduces communication costs in distributed learning by 95%.
Model compression techniques, such as pruning and quantization, are becoming increasingly important to reduce the memory footprints and the amount of computations. Despite model size reduction, achieving performance enhancement on devices is, however, still challenging mainly due to the irregular representations of spa…
Accelerated magnetic resonance (MR) scan acquisition with compressed sensing (CS) and parallel imaging is a powerful method to reduce MR imaging scan time. However, many reconstruction algorithms have high computational costs. To address this, we investigate deep residual learning networks to remove aliasing artifacts …
Tensor decompositions are powerful tools for large data analytics as they jointly model multiple aspects of data into one framework and enable the discovery of the latent structures and higher-order correlations within the data. One of the most widely studied and used decompositions, especially in data mining and machi…
Improved SGD rates with delayed and compressed gradients.
This the first of a set of three papers about the Compression Theorem: if M^m is embedded in Q^q X R with a normal vector field and if q-m > 0, then the given vector field can be straightened (ie, made parallel to the given R direction) by an isotopy of M and normal field in Q X R. The theorem can be deduced from Gromo…
As a result of the growing size of Deep Neural Networks (DNNs), the gap to hardware capabilities in terms of memory and compute increases. To effectively compress DNNs, quantization and connection pruning are usually considered. However, unconstrained pruning usually leads to unstructured parallelism, which maps poorly…
New methods accelerate NCGP inference by trading computation for uncertainty.
This paper combines three techniques to reduce communications in distributed variational inequalities.
This monograph presents a class of algorithms called coordinate descent algorithms for mathematicians, statisticians, and engineers outside the field of optimization. This particular class of algorithms has recently gained popularity due to their effectiveness in solving large-scale optimization problems in machine lea…
This is the third of three papers about the Compression Theorem: if M^m is embedded in Q^q X R with a normal vector field and if q-m > 0, then the given vector field can be straightened (ie, made parallel to the given R direction) by an isotopy of M and normal field in Q X R. The theorem can be deduced from Gromov's th…
New algorithms allow multiple robots to search efficiently without central coordination.
A new method compresses NLP networks by using multiple subspaces instead of a single one.
We describe the multi-GPU gradient boosting algorithm implemented in the XGBoost library (https://github.com/dmlc/xgboost). Our algorithm allows fast, scalable training on multi-GPU systems with all of the features of the XGBoost library. We employ data compression techniques to minimise the usage of scarce GPU memory …
A practical limitation of deep neural networks is their high degree of specialization to a single task and visual domain. Recently, inspired by the successes of transfer learning, several authors have proposed to learn instead universal, fixed feature extractors that, used as the first stage of any deep network, work w…
We show that bordered Heegaard Floer homology detects incompressible surfaces and bordered-sutured Floer homology detects partly boundary parallel tangles and bridges, in natural ways. For example, there is a bimodule Lambda so that the tensor product of CFD(Y) and Lambda is Hom-orthogonal to CFD(Y) if and only if the …
Suppose N is a compressible boundary component of a compact orientable irreducible 3-manifold M and Q is an orientable properly embedded essential surface in M in which each component is incident to N and no component is a disk. Let VN and QN denote respectively the sets of vertices in the curve complex for N represent…
Large-scale L1-regularized loss minimization problems arise in high-dimensional applications such as compressed sensing and high-dimensional supervised learning, including classification and regression problems. High-performance algorithms and implementations are critical to efficiently solving these problems. Building…
Source coding is the canonical problem of data compression in information theory. In a locally encodable source coding, each compressed bit depends on only few bits of the input. In this paper, we show that a recently popular model of semi-supervised clustering is equivalent to locally encodable source coding. In this …