A new quantization strategy reduces Transformer model size and inference time.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper explores the problem of learning transforms for image compression via autoencoders. Usually, the rate-distortion performances of image compression are tuned by varying the quantization step size. In the case of autoen-coders, this in principle would require learning one transform per rate-distortion point at…
BiTAT improves neural network quantization for edge devices by focusing on weight dependencies and disentangling them.
State-of-the-art neural machine translation methods employ massive amounts of parameters. Drastically reducing computational costs of such methods without affecting performance has been up to this point unsuccessful. To this end, we propose FullyQT: an all-inclusive quantization strategy for the Transformer. To the bes…
DFRot improves LLMs by reducing outlier and massive activation effects.
The paper studies quantization on symplectic manifolds with real polarizations, comparing different quantization methods.
Smooth approximations of Kähler-Ricci solitons found using quantized metrics and Futaki invariants.
In this lecture results on the Berezin-Toeplitz quantization of arbitrary compact quantizable Kaehler manifolds are presented. These results are obtained in joint work with M. Bordemann and E. Meinrenken. The existence of the Berezin-Toeplitz deformation quantization is also covered. Recent results obtained in joint wo…
New method preserves spectral clustering performance under aggressive sparsification and quantization.
Construct geometric interpretation of Heston model using group quantization.
Differentiable model compression adds noise to parameters during training.
A new method quantizes output space for multi-target regression.
We extend quantization-aware training to extreme model compression.
It is shown that the heat operator in the Hall coherent state transform for a compact Lie group is related with a Hermitian connection associated to a natural one-parameter family of complex structures on . The unitary parallel transport of this connection establishes the equivalence of (geometric) quantizati…
Mackey showed that for a compact Lie group , the pair has a unique non-trivial irreducible covariant pair of representations. We study the relevance of this result to the unitary equivalence of quantizations for an infinite-dimensional family of invariant polarizations on . The …
S2D selectively decays large singular values to improve quantization of neural activations.
Quantizes contact structures using dynamical methods.
Quantizes symplectic fibrations to analyze vector bundles and metrics.
We present Rotated Adaptive Tetra-iterated Quantizer (RATQ), a fixed-length quantizer for gradients in first order stochastic optimization. RATQ is easy to implement and involves only a Hadamard transform computation and adaptive uniform quantization with appropriately chosen dynamic ranges. For noisy gradients with al…
For phase-space manifolds which are compact Kaehler manifolds relations between the Berezin-Toeplitz quantization and the quantization with the help of Berezin's coherent states and symbols are studied. First the results on the Berezin-Toeplitz quantization of arbitrary compact Kaehler manifolds due to Bordemann, Meinr…
A method for vectorizing persistence diagrams simplifies topological data analysis.
New method shows unitarity in quantization for toric manifolds.
The paper is devoted to the comparison of the Fukaya category (it is responcible for the A-side of mirror symmetry) with the category of holonomic modules over the quantized algebra of functions on the same symplectic manifold. We conjecture that these categories become -equivalent after a twist by a kind o…
The Cauchy problem for the homogeneous (real and complex) Monge-Ampere equation (HRMA/HCMA) arises from the initial value problem for geodesics in the space of Kahler metrics. It is an ill-posed problem. We conjecture that, in its lifespan, the solution can be obtained by Toeplitz quantizing the Hamiltonian flow define…
New insights into quantized neural networks reveal learning dynamics and generalization errors.
Neural network quantization has become an important research area due to its great impact on deployment of large models on resource constrained devices. In order to train networks that can be effectively discretized without loss of performance, we introduce a differentiable quantization procedure. Differentiability can…
In the first part of the paper we describe the complex geometry of the universal Teichmüller space , which may be realized as an open subset in the complex Banach space of holomorphic quadratic differentials in the unit disc. The quotient of the diffeomorphism group of the circle modulo Möbius …
This paper provides theoretical foundations for using quantized actions in behavior cloning.
A novel method quantizes Batch Normalization for QNNs, maintaining accuracy and efficiency.
Let be a pseudo-Riemannian manifold and the space of densities of degree on . We study the space of second-order differential operators from to . If is conformally flat with signature , then is viewed as a module over the group of confo…
Mathematical framework for brane quantization using SYZ mirror symmetry.
For the cotangent bundle of a compact Lie group , we study the complex-time evolution of the vertical tangent bundle and the associated geometric quantization Hilbert space under an infinite-dimensional family of Hamiltonian flows. For each such flow, we construct a generalized coherent state tra…
For any Lie groupoid , the vector bundle dual to the associated Lie algebroid is canonically a Poisson manifold. The (reduced) C*-algebra of (as defined by A. Connes) is shown to be a strict quantization (in the sense of M. Rieffel) of . This is proved using a generalization of Weyl's quantization…
Study develops numerical schemes for non-Markovian volatility models with memory.
This paper proposes a new subspace learning method, named Quantized Fisher Discriminant Analysis (QFDA), which makes use of both machine learning and information theory. There is a lack of literature for combination of machine learning and information theory and this paper tries to tackle this gap. QFDA finds a subspac…
Paper proposes a method to speed up DNNs by quantizing Winograd/Toom-Cook convolutions.
Transformers can cluster data from Gaussian mixtures without supervision.
An index formula is proposed for contact transformations between contact manifolds equipped with CR structures or with fillings by symplectic manifolds. The formula generalizes the Atiyah-Singer formula and gives a conjectured formula for the index of Fourier integral operators, as well as Epstein's relative index for …
Lightweight architectural designs of Convolutional Neural Networks (CNNs) together with quantization have paved the way for the deployment of demanding computer vision applications on mobile devices. Parallel to this, alternative formulations to the convolution operation such as FFT, Strassen and Winograd, have been ad…
A membrane technique, in which the symplectic and Ricci forms are integrated over surfaces in a complexification of the phase space, as well a ``creation" connection with zero curvature over lagrangian submanifolds, is used to obtain a unified quantization including a noncommutative algebra of functions, its representa…
This paper develops quantization algorithms for random Fourier features, simplifying the process and improving performance.
We quantize the interaction of gravity with Yang-Mills and spinor fields, hence offering a quantum theory incorporating all four fundamental forces of nature. Using canonical quantization we obtain solutions of the Wheeler-DeWitt equation in a vector bundle and the method of second quantization leads to a symplectic ve…
Efficiently processes dynamic inputs in AI writing assistants with incremental computation.
TimeVQVAE uses VQ for better time series generation.
Theorems adapted for prequantum systems, showing symplectomorphism and gauge transformation.
Infinitesimal conformal transformations of are always polynomial and finitely generated when . Here we prove that the Lie algebra of infinitesimal conformal polynomial transformations over , , is maximal in the Lie algebra of polynomial vector fields. When is greater than 2 and are such t…
Optimal quantization improves dataset distillation for faster training.
In this paper, we study the analytic continuation to complex time of the Hamiltonian flow of certain -invariant functions on the cotangent bundle of a compact connected Lie group with maximal torus . Namely, we will take the Hamiltonian flows of one -invariant function, , and one $G\time…