Improved machine translation with INT8 hardware using a novel training method.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
FrostNet improves INT8 quantization efficiency in mobile networks.
Efficient Winograd convolution for INT8 networks using RNS.
Study finds optimal learning rate schedules for sub-100M quantization-aware training across bit-widths.
MSD removes dequantization bottleneck in LLM inference by approximating high-precision activations.
Degree-Quant improves GNN efficiency by quantizing them without losing accuracy.
This study improves Ernie's accuracy for INT8 inference by modifying its training process.
Generative compression technique reduces neural network size and improves performance on microcontrollers.
Paper develops a statistical framework for quantized training of deep neural networks.
We extend quantization-aware training to extreme model compression.
APQ jointly optimizes neural architecture, pruning, and quantization for efficient inference.
RFX accelerates and compresses Random Forests for large datasets.