APoT quantization improves neural network efficiency and accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method achieves faster calibration without randomization.
Efficiently detects anomalies in videos with reduced computation.
Let K be a connected finite complex. This paper studies the problem of whether one can attach a cell to some iterated suspension S^j K so that the resulting space satisfies Poincare duality. When this is possible, we say that S^j K is a spine. We introduce the notion of quadratic self duality and show that if K is quad…
We show that the only rational homology spheres which can admit almost complex structures occur in dimensions two and six. Moreover, we provide infinitely many examples of six-dimensional rational homology spheres which admit almost complex structures, and infinitely many which do not. We then show that if a closed alm…
The use of low-precision fixed-point arithmetic along with stochastic rounding has been proposed as a promising alternative to the commonly used 32-bit floating point arithmetic to enhance training neural networks training in terms of performance and energy efficiency. In the first part of this paper, the behaviour of …
In this paper, we study how the mean shift algorithm can be used to denoise a dataset. We introduce a new framework to analyze the mean shift algorithm as a denoising approach by viewing the algorithm as an operator on a distribution function. We investigate how the mean shift algorithm changes the distribution and sho…
Randomly initialized ReLU networks of depth two can approximate smooth functions well.
Let Gamma be a group generated by two positive multi-twists. We give some sufficient conditions for Gamma to be free or have no `unexpectedly reducible' elements. For a group Gamma generated by two Dehn twists, we classify the elements in Gamma which are multi-twists. As a consequence we are able to list all the lanter…
Alternatives to recurrent neural networks, in particular, architectures based on attention or convolutions, have been gaining momentum for processing input sequences. In spite of their relevance, the computational properties of these alternatives have not yet been fully explored. We study the computational power of two…
We consider derivative-free algorithms for stochastic and non-stochastic convex optimization problems that use only function values rather than gradients. Focusing on non-asymptotic bounds on convergence rates, we show that if pairs of function values are available, algorithms for -dimensional optimization that use …
Operating deep neural networks on devices with limited resources requires the reduction of their memory footprints and computational requirements. In this paper we introduce a training method, called look-up table quantization, LUT-Q, which learns a dictionary and assigns each weight to one of the dictionary's values. …
HMQ improves quantization for edge devices with mixed precision.
For any Legendrian knot in standard contact we relate counts of ungraded (-graded) representations of the Legendrian contact homology DG-algebra with the -colored Kauffman polynomial. To do this, we introduce an ungraded -colored ruling polynomial, …
Cyber-Physical Systems (CPSs) have been pervasive including smart grid, autonomous automobile systems, medical monitoring, process control systems, robotics systems, and automatic pilot avionics. As usually implemented on embedded devices, CPS is typically constrained by computation capacity and energy consumption. In …
Generative adversarial networks (GANs) are innovative techniques for learning generative models of complex data distributions from samples. Despite remarkable recent improvements in generating realistic images, one of their major shortcomings is the fact that in practice, they tend to produce samples with little divers…
The use of deep neural networks in edge computing devices hinges on the balance between accuracy and complexity of computations. Ternary Connect (TC) \cite{lin2015neural} addresses this issue by restricting the parameters to three levels , and , thus eliminating multiplications in the forward pass of the net…
DJPQ optimizes neural network pruning and quantization for hardware efficiency.
Bayesian Bits unifies quantization and pruning through gradient optimization.
We consider the problem of deep neural net compression by quantization: given a large, reference net, we want to quantize its real-valued weights using a codebook with entries so that the training loss of the quantized net is minimal. The codebook can be optimally learned jointly with the net, or fixed, as for bina…
Paper details Hilbert-curve for high-performance data mining.
MAGDiff detects data shifts in neural networks without retraining.
Deep convolutional neural networks (CNNs) are powerful tools for a wide range of vision tasks, but the enormous amount of memory and compute resources required by CNNs pose a challenge in deploying them on constrained devices. Existing compression techniques, while excelling at reducing model sizes, struggle to be comp…
A new iterative low complexity algorithm has been presented for computing the Walsh-Hadamard transform (WHT) of an dimensional signal with a -sparse WHT, where is a power of two and , scales sub-linearly in for some . Assuming a random support model for the non-zero transform domain…
Modern deep learning models are often trained in parallel over a collection of distributed machines to reduce training time. In such settings, communication of model updates among machines becomes a significant performance bottleneck and various lossy update compression techniques have been proposed to alleviate this p…
RED-2400 is a public benchmark of trading events from a Solana exchange, labeled by algorithmic rejection.
Computes volumes of metric maps on surfaces, linking to Weil-Petersson volumes.
This paper proposes a learning framework for n-bit quantized neural networks that improves accuracy and speed on FPGAs.