Optimizes tensor rank selection for neural network compression.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method compresses neural networks up to 14x with minimal performance loss.
Bayesian neural networks are compressed using feature and weight pruning based on posterior inclusion probabilities.
Convolutional neural networks (CNNs) achieve state-of-the-art performance in a wide variety of tasks in computer vision. However, interpreting CNNs still remains a challenge. This is mainly due to the large number of parameters in these networks. Here, we investigate the role of compression and particularly pruning fil…
Physics-inspired methods optimize SVD compression of LLMs.
Compressing word embeddings is important for deploying NLP models in memory-constrained settings. However, understanding what makes compressed embeddings perform well on downstream tasks is challenging---existing measures of compression quality often fail to distinguish between embeddings that perform well and those th…
Auto-Compressing Subset Pruning reduces model size for faster inference.
Improves convergence speed in compressive sensing with a new probabilistic approach.
A new method compresses conditional distributions of labelled data.
This work optimizes DNN inference for energy-harvesting devices by compressing and selectively executing neural network exits.
End-to-end automatic speech recognition (ASR) models are increasingly large and complex to achieve the best possible accuracy. In this paper, we build an AutoML system that uses reinforcement learning (RL) to optimize the per-layer compression ratios when applied to a state-of-the-art attention based end-to-end ASR mod…
The low-rank tensor approximation is very promising for the compression of deep neural networks. We propose a new simple and efficient iterative approach, which alternates low-rank factorization with a smart rank selection and fine-tuning. We demonstrate the efficiency of our method comparing to non-iterative ones. Our…
Optimizes MCMC output compression by selecting a subset of states.
Iterative AutoML improves ASR model compression by 5x without WER degradation.
HOTCAKE compresses CNNs by decomposing kernels into smaller parts.
MARS automatically selects tensor decomposition ranks, improving performance in neural network tasks.
Improved survival analysis using square root Cox's models and neural networks.
A new method for high-dimensional data classification reduces misclassification errors.
We present a new similarity measure based on information theoretic measures which is superior than Normalized Compression Distance for clustering problems and inherits the useful properties of conditional Kolmogorov complexity. We show that Normalized Compression Dictionary Size and Normalized Compression Dictionary En…
Galen algorithm compresses neural networks for specific hardware with reduced latency.
Three new efficient algorithms project vectors onto weighted l1 ball.
Deep learning accelerators efficiently train over vast and growing amounts of data, placing a newfound burden on commodity networks and storage devices. A common approach to conserve bandwidth involves resizing or compressing data prior to training. We introduce Progressive Compressed Records (PCRs), a data format that…
We present reconstruction algorithms for smooth signals with block sparsity from their compressed measurements. We tackle the issue of varying group size via group-sparse least absolute shrinkage selection operator (LASSO) as well as via latent group LASSO regularizations. We achieve smoothness in the signal via fusion…
CoDeQ simplifies joint model compression by integrating pruning and quantization.
New SVM margin bound improves generalization in machine learning.
This paper introduces ASCAI, a novel adaptive sampling methodology that can learn how to effectively compress Deep Neural Networks (DNNs) for accelerated inference on resource-constrained platforms. Modern DNN compression techniques comprise various hyperparameters that require per-layer customization to ensure high ac…
We introduce and study the problem of Online Continual Compression, where one attempts to simultaneously learn to compress and store a representative dataset from a non i.i.d data stream, while only observing each sample once. A naive application of auto-encoders in this setting encounters a major challenge: representa…
Proposes a method to improve Byzantine-robustness in compressed federated learning.
We investigate an optimal investment problem with a general performance criterion which, in particular, includes discontinuous functions. Prices are modeled as diffusions and the market is incomplete. We find an explicit solution for the case of limited diversification of the portfolio, i.e. for the portfolio compressi…
Optimized sampling scheme for compressed sensing combining randomness and determinism.
Proposes a new method to selectively access privileged information in reinforcement learning.
This paper presents a novel signal compression algorithm based on the Blaschke unwinding adaptive Fourier decomposition (AFD). The Blaschke unwinding AFD is a newly developed signal decomposition theory. It utilizes the Nevanlinna factorization and the maximal selection principle in each decomposition step, and achieve…
A new method enhances signal recovery with FDR control.
Modality-agnostic compression improves across diverse data types.
Highly distributed training of Deep Neural Networks (DNNs) on future compute platforms (offering 100 of TeraOps/s of computational capacity) is expected to be severely communication constrained. To overcome this limitation, new gradient compression techniques are needed that are computationally friendly, applicable to …
Paper compresses deep neural networks by eliminating redundant neurons.
Compression techniques for deep neural networks are important for implementing them on small embedded devices. In particular, channel-pruning is a useful technique for realizing compact networks. However, many conventional methods require manual setting of compression ratios in each layer. It is difficult to analyze th…
Clustering, like covariate selection for classification, is an important step to compress and interpret the data. However, clustering of covariates is often performed independently of the classification step, which can lead to undesirable clustering results that harm interpretability and compression rate. Therefore, we…
End-to-end meta-learned system for image compression.
In recent years, deep neural networks have achieved great success in the field of computer vision. However, it is still a big challenge to deploy these deep models on resource-constrained embedded devices such as mobile robots, smart phones and so on. Therefore, network compression for such platforms is a reasonable so…
New method compresses deep learning layers using tensor decomposition.
New method optimizes lossy compression models more effectively.
We frame the problem of selecting an optimal audio encoding scheme as a supervised learning task. Through uniform convergence theory, we guarantee approximately optimal codec selection while controlling for selection bias. We present rigorous statistical guarantees for the codec selection problem that hold for arbitrar…
SSVI efficiently trains sparse Bayesian neural networks with minimal compression and performance loss.
New approach selects sparse features without validation.
Tensor decomposition is an effective approach to compress over-parameterized neural networks and to enable their deployment on resource-constrained hardware platforms. However, directly applying tensor compression in the training process is a challenging task due to the difficulty of choosing a proper tensor rank. In o…
We propose a modification of linear discriminant analysis, referred to as compressive regularized discriminant analysis (CRDA), for analysis of high-dimensional datasets. CRDA is specially designed for feature elimination purpose and can be used as gene selection method in microarray studies. CRDA lends ideas from $\el…
Paper uses neural networks to compress large portfolios of options, reducing risk and capital requirements.