Paper proposes a 1-bit quantization scheme for high-dimensional statistical estimation.
problem High-dimensional statistical estimation with limited data.
method Uniformly dithered 1-bit quantization for sparse covariance matrix estimation, sparse linear regression, and matrix completion.
result Near minimax rates in sub-Gaussian regime and improved rates in heavy-tailed regime.
A new method for quantized matrix completion using Huber loss.
problem Quantized Matrix Completion with robustness to quantization errors.
method Rank minimization with Huber loss regularization, Smooth Rank Approximation.
result Our method achieves better accuracy and efficiency than state-of-the-art methods.
Paper quantizes heavy-tailed data for near optimal estimation rates.
problem Estimating parameters from heavy-tailed data with quantization.
method Truncate and dither data, then uniformly quantize; achieves near minimax rates.
result Near optimal estimation rates achievable with quantized data.
The task of estimating a matrix given a sample of observed entries is known as the \emph{matrix completion problem}. Most works on matrix completion have focused on recovering an unknown real-valued low-rank matrix from a random sample of its entries. Here, we investigate the case of highly quantized observations when …
New method corrects quantization errors in LLMs using low-rank matrices.
problem Correcting quantization errors in large language models.
method Introducing low-rank weight matrices to correct quantized activations in LLMs.
result Reduces accuracy gap with original model by more than 50% using low-rank matrices.
QuantEase optimizes LLMs with CD-based quantization, achieving state-of-the-art performance.
problem Efficiently quantize large language models for deployment.
method Layer-wise quantization using CD-based algorithms with matrix and vector operations.
result State-of-the-art performance in perplexity and zero-shot accuracy.
Abstract proposes a new categorical approach to quantization of Poisson algebras.
problem Quantization of Poisson algebras.
method Defining quantization categories as subcategories of R-module categories with classical limits.
result Categories of strict deformation quantization, prequantization, and matrix regularization are equivalent, while Poisson enveloping algebra is not.
The paper analyzes how quantization affects the Fisher Information Matrix's dominant eigenvalue.
problem The impact of quantization on the Fisher Information Matrix's dominant eigenvalue.
method The study examines spectral perturbation of the empirical Fisher Information Matrix under in-distribution input and quantized parameter perturbations.
result A bound on the eigenvalue under quantization noise, showing it strictly exceeds the unperturbed value at leading order.
Survey on matrix hydrodynamics, a 2D fluid model.
problem Modeling 2D incompressible fluids.
method Spatial discretization via quantization theory.
result Basic demonstration of matrix hydrodynamics.
BiQGEMM efficiently multiplies quantized DNN weights using lookup tables.
problem Efficiently multiplying quantized DNN weights on CPUs/GPUs with limited memory.
method BiQGEMM pre-computes and stores redundant intermediate results in lookup tables.
result BiQGEMM achieves lower overall computations and higher performance.
New method preserves spectral clustering performance under aggressive sparsification and quantization.
problem Maintaining spectral clustering performance with sparse and quantized data.
method Random matrix theory applied to eigenspectrum changes under sparsification and quantization.
result Spectral clustering performance is preserved even with aggressive sparsification and quantization.
Optimization problems with rank constraints arise in many applications, including matrix regression, structured PCA, matrix completion and matrix decomposition problems. An attractive heuristic for solving such problems is to factorize the low-rank matrix, and to run projected gradient descent on the nonconvex factoriz…
A matrix completion problem, which aims to recover a complete matrix from its partial observations, is one of the important problems in the machine learning field and has been studied actively. However, there is a discrepancy between the mainstream problem setting, which assumes continuous-valued observations, and some…
In this paper, we consider the matrix completion problem when the observations are one-bit measurements of some underlying matrix M, and in particular the observed samples consist only of ones and no zeros. This problem is motivated by modern applications such as recommender systems and social networks where only "like…
Paper studies quantized LRMR with random dithering for correlated tasks.
problem Estimating coefficient matrix in quantized multivariate regression.
method Uniform quantization with random dithering, constrained and regularized Lasso estimators.
result Achieves minimax optimal rate with dithering, slightly worsens quantization effect.
11D supergravity completes with quantized C-field flux.
problem Completing 11D supergravity with quantized C-field flux.
method Duality-symmetric formulation of on-shell 11d supergravity on superspace.
result 11d super-spacetimes are quantizable by duality-symmetric super-C-field flux.
Study asymptotics of unitary matrix elements in quantum mechanics.
problem Asymptotic behavior of unitary matrix elements in quantum mechanics.
method Uses Berezin-Toeplitz quantization and symplectic geometry.
result Recover asymptotics of Wigner's d-matrix elements for spin representations.
This paper examines fundamental error characteristics for a general class of matrix completion problems, where the matrix of interest is a product of two a priori unknown matrices, one of which is sparse, and the observations are noisy. Our main contributions come in the form of minimax lower bounds for the expected pe…
This paper examines a general class of noisy matrix completion tasks where the goal is to estimate a matrix from observations obtained at a subset of its entries, each of which is subject to random noise or corruption. Our specific focus is on settings where the matrix to be estimated is well-approximated by a product …
DFRot improves LLMs by reducing outlier and massive activation effects.
problem Reducing outlier and massive activation effects in rotated LLMs.
method Weighted loss function and orthogonal Procrustes transforms for rotation matrix refinement.
result DFRot achieves dual free (Outlier-Free and Massive Activation-Free) with significant improvements in perplexity.
It is postulated that quantum gravity is a sum over causal structures coupled to matter via scale evolution. Quantized causal structures can be described by studying simple matrix models where matrices are replaced by an algebra of quantum mechanical observables. In particular, previous studies constructed quantum grav…
I exhibit a prequantization of the torus which is actually a ``full'' quantization in the sense that a certain complete set of classical observables is irreducibly represented. Thus in this instance there is no Groenewold-Van Hove obstruction to quantization.
This work proposes a complete 8-bit quantization framework for large-scale deep neural networks.
problem Training large-scale deep neural networks with high performance and low memory footprint.
method WAGEUBN framework that quantizes all data paths including weights, activations, gradients, errors, updates, and batch normalization.
result Achieves competitive accuracy on the ImageNet dataset using only 8-bit integers.
In this paper we pursue the study of formal geometric quantization of non-compact Hamiltonian manifolds. Our main result is the proof that two quantization process coincide. This fact was obtained by Ma and Zhang in the preprint arXiv:0812.3989 by completely different means.
Motivated by the problem of deformation quantization we introduce and study directed graph complexes with oriented loops and wheels. We develop some technique for computing cohomology of such graph complexes and apply it to several concrete examples such as wheeled completion of the operad of strongly homotopy Lie alge…
The recently proposed SPARse Factor Analysis (SPARFA) framework for personalized learning performs factor analysis on ordinal or binary-valued (e.g., correct/incorrect) graded learner responses to questions. The underlying factors are termed "concepts" (or knowledge components) and are used for learning analytics (LA),…
We determine the matrix of the balanced metric of the Siegel-Jacobi ball and its inverse. We calculate the scalar curvature, the Ricci form and the Laplace-Beltrami operator of this manifold. We discuss several geometric aspects related with Berezin quantization on the Siegel-Jacobi ball.
In the framework of geometric quantization we extend the Bohr-Sommerfeld rules to a full quantization theory which resembles Heisenberg's matrix theory. This extension is possible because Bohr-Sommerfeld rules not only provide an orthogonal basis in the space of quantum states, but also give a lattice structure to this…
Deep task-based quantization improves MIMO signal processing.
problem Improving performance of MIMO signal processing with scalar ADCs.
method Data-driven task-oriented quantization using deep learning.
result Deep task-based quantization can approach optimal performance limits.
Modernizes higher-dimensional supergravity, linking it to flux quantization.
problem Constructing infrared completions of higher-dimensional supergravity.
method Using differential nonabelian cohomology and super-torsion constraints.
result Equivalence of solutions in different dimensions and flux quantization.
Any classical r-matrix on the Lie algebra of linear operators on a real vector space V gives rise to a quadratic Poisson structure on V which admits a deformation quantization stemming from the construction of V. Drinfel'd. We exhibit in this article an example of quadratic Poisson structure which does not arise this w…
Defines coherent manifolds and their quantum applications.
problem Understanding quantum spaces and coherent states.
method Generalizes quantum field theory and Schrödinger equation solutions.
result Solves Schrödinger equation on coherent manifolds.
Paper proves existence of a universal codebook for low-precision quantization.
problem Optimizing low-precision approximation of matrix products in machine learning.
method Develops a universal codebook that is near-optimal for all possible statistics of input data.
result Proves existence of a universal codebook with a 0.11 bit per dimension reduction in rate.
The study quantizes ancient flows in cylinders, revealing their asymptotic behavior.
problem Analyzing ancient mean curvature flows with cylindrical tangent profiles.
method Proved asymptotic behavior of cylindrical profile functions using spectral quantization.
result Asymptotic behavior of cylindrical profile functions quantized to eigenvalues 0 or -sqrt(2(n-k))/4.
In this paper we continue our study of Groenewold-Van Hove obstructions to quantization. We show that there exists such an obstruction to quantizing the cylinder T∗S1. More precisely, we prove that there is no quantization of the Poisson algebra of T∗S1 which is irreducible on a naturally defined $e(2) \times R…
Improved 2-bit covariance estimator with reduced operator norm error and no tuning needed.
problem Improving 2-bit covariance estimation with reduced operator norm error and no tuning needed.
method Proposed a new 2-bit covariance matrix estimator using triangular dithering scales.
result Improved operator norm error rate that depends on effective rank of covariance matrix, closing theoretical gap.
LVQ models robustness evaluated against adversarial attacks.
problem Robustness of LVQ models against adversarial attacks.
method Evaluation of three LVQ models: Generalized LVQ, Generalized Matrix LVQ, and Generalized Tangent LVQ.
result Generalized LVQ and Generalized Tangent LVQ are robust, while Generalized Matrix LVQ is not.
Algorithm compresses large matrices by approximating them as low rank and low precision factors.
problem Efficiently storing and processing large matrices with billions of elements.
method Randomized sketching and quantization of matrix columns to achieve low rank and low precision factorization.
result Achieves compression ratios as low as one bit per matrix coordinate while maintaining or improving performance.
A new method improves maximum inner product search by locally decomposing residual vectors.
problem Maximum inner product search efficiency and accuracy.
method Local Orthogonal Decomposition (LOD) combined with multiscale quantization.
result LOD consistently achieves higher recall than previous methods under the same bitrates.
Paper shows how to integrate quantization into neural compression models.
problem Integrating quantization into neural compression models.
method Integrates uniform noise channel at test time using universal quantization.
result Eliminates mismatch between training and test phases while maintaining differentiability.
This paper finds a new way to compress CNN weights, improving on pruning and quantization.
problem Improving performance and storage efficiency of CNNs.
method Identifying and exploiting repeated patterns in CNN weight tensors, using Huffman coding and block sparse matrix formats.
result Achieved compaction ratios of 1.4x to 3.1x in addition to pruning and quantization.
This paper tackles tensor recovery from noisy and multi-level quantized measurements.
problem Tensors from multi-level quantized measurements.
method Nonconvex optimization problem with alternating proximal gradient descent.
result The recovery error diminishes to zero with increasing tensor dimensions.
Paper proposes method to recover quantized data with missing info.
problem Recovering quantized data with missing information.
method Regularized convex cost function with Bi-factorization and Augmented Lagrangian Method.
result The method finds global minimizer of the cost function.
Study integrability of quantized six-vertex model on torus.
problem Integrability of a specific lattice model on a torus.
method Defined layer transfer matrices and tetrahedron equations for admissible graphs.
result Established commutativity of transfer matrices and derived quantum Hamiltonians.
BiTAT improves neural network quantization for edge devices by focusing on weight dependencies and disentangling them.
problem Performance degradation of compact neural networks under extreme quantization.
method Task-dependent Aggregated Transformation (BiTAT) method that orthonormalizes weights and progressively quantizes them.
result BiTAT effectively preserves model performance on ImageNet and CIFAR-100 with compact backbones.
A method for reducing neural network size using look-up tables.
problem Reducing memory and computational footprint of deep neural networks.
method Iteratively learns value dictionaries and assignment matrices for network weights.
result General framework for network reduction that can handle various reduction problems.
Single model corrects JPEG artifacts for various compression settings.
problem JPEG compression artifacts due to aggressive quantization.
method Parameterized architecture using quantization matrix.
result State-of-the-art performance across different quality settings.
M5-branes' flux quantization linked to non-abelian cohomology.
problem Flux quantization on M5-branes and its implications.
method Analogous to Dirac's charge/flux quantization, constraining M5's flux-quantization law to non-abelian cohomology theory.
result Skyrmion-like and anyonic solitons on M5-branes and open M5-branes.