Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

106212317423 · Jun 202019922001200920182026
48 results for floating point reduction

Binarized CNNs improve GPU inference efficiency on resource-constrained devices.

problem Efficient inference on resource-constrained devices for image classification.
method Binarization of weights and computations in CNNs, implemented on GPUs.
result 7.4X speedup with 4.4% accuracy loss on embedded GPU platforms.

No floating point, no multiplications, no problem! Training efficient networks for resource-constrained devices.

problem Designing efficient neural networks for resource-constrained devices without floating-point operations.
method Discretizing both in-network non-linearities and network weights to avoid floating-point and multiplication operations.
result Training networks without floating-point operations can achieve comparable performance to those using floating-point operations, with less memory usage.

Researchers find floating point errors can mislead neural network verifiers.

problem Floating point arithmetic inaccuracies mislead neural network verifiers.
method Efficiently searches inputs and constructs neural network architectures to exploit verification errors.
result Floating point errors can systematically mislead neural network verifiers.

Paper analyzes Winograd convolution errors and proposes methods to reduce them.

problem Reduction of floating point error in Winograd convolution for deep neural networks.
method Analysis of worst case FP error, estimation of norm and conditioning, proposed evaluation orderings, sampling points selection, mixed-precision convolution, pairwise summation.
result Proposed methods significantly reduce FP error for a given block size, allowing larger block sizes and reduced computation.

Channel gating reduces CNN computation cost by skipping ineffective feature regions.

problem Reducing computation cost in CNNs while maintaining accuracy.
method Dynamic, fine-grained pruning scheme that identifies and skips computation on ineffective feature regions.
result 2.7-8.0x reduction in FLOPs and 2.0-4.4x reduction in memory accesses with minimal accuracy loss.

Paper explores low-precision arithmetic for neural networks, improving efficiency.

problem Training neural networks with high precision and energy efficiency.
method 12-bit fixed-point, 12-bit floating-point, local scaling, Power-of-Two arithmetic.
result 7-bit Power-of-Two arithmetic achieves minimal loss in accuracy with reduced computation.

Training deep neural networks with 8-bit floating point numbers is now possible and more efficient.

problem Challenges in training DNNs with reduced precision, especially for gradient computations.
method Introduction of chunk-based accumulation and floating point stochastic rounding to reduce arithmetic precision to 16 bits.
result Successful training of DNNs using 8-bit floating point numbers, maintaining accuracy on various models and datasets.

Paper proposes training deep neural networks with 8-bit floating point precision.

problem Challenges in training deep neural networks at 8-bit precision due to higher precision and dynamic range requirements.
method Proposes a method to train deep neural networks using 8-bit floating point for weights, activations, errors, and gradients. Introduces an enhanced loss scaling method and stochastic rounding technique.
result Demonstrates state-of-the-art accuracy across multiple datasets and workloads compared to full precision baseline.

A method to reduce knowledge graph embedding models by binarizing parameters.

problem Large memory requirements for tensor factorization models in knowledge graph completion.
method Introducing a quantization function to binarize parameters of CP tensor decomposition.
result Successfully reduced model size by more than an order of magnitude while maintaining task performance.

Review and compare model order reduction methods for process engineering.

problem Creating computationally efficient yet accurate models for real-time applications.
method Nonlinear model order reduction methods, including general-purpose and tailored approaches for chemical processes.
result Comparison of eight model order reduction methods applied to an air separation process model.

We present a theoretical analysis and empirical evaluations of a novel set of techniques for computational cost reduction of classifiers that are based on learned transform and soft-threshold. By modifying optimization procedures for dictionary and classifier training, as well as the resulting dictionary entries, our t…

2015-04-26abs ↗pdf ↗

Cheetah framework optimizes DNNs for edge devices using low-precision formats.

problem Reducing DNN model size for edge devices while maintaining accuracy.
method Mixed low-precision hardware and software co-design framework using posit and other formats.
result 16-bit posits outperform 16-bit floating point in training, and [5..8]-bit posits improve inference performance.

Custom narrow-precision representations boost DNN inference speed by 7.6x with minimal accuracy loss.

problem Improving computational efficiency of deep neural networks.
method Exploring and utilizing unconventional narrow-precision floating-point representations for DNN weights and activations.
result Average speedup of 7.6x with less than 1% accuracy loss.

StatQAT optimizes quantization for deep networks, reducing computational cost and memory usage.

problem Optimal quantization parameters selection for deep neural networks with diverse data distributions.
method Statistical error analysis framework for uniform and floating-point quantization, iterative and analytic quantizers designed for arbitrary and Gaussian-like distributions.
result Improved accuracy and stability in training low-precision neural networks.

Quantization scheme improves inference efficiency for deep learning models.

problem Limited computational resources and energy constraints in deploying deep learning models.
method Proposes a quantization scheme to reduce precision and improve efficiency without sacrificing accuracy.
result End-to-end post-quantization accuracies comparable to reference model achieved using a single inference batch calibration.

For a convex body on the Euclidean unit sphere the spherical convex floating body is introduced. The asymptotic behavior of the volume difference of a spherical convex body and its spherical floating body is investigated. This gives rise to a new spherical area measure, the floating area. Remarkably, this floating area…

2014-11-27abs ↗pdf ↗

Flexpoint improves deep learning training efficiency by using adaptive 16-bit format.

problem Training deep neural networks in low bit-width formats is challenging.
method Flexpoint uses a shared exponent dynamically adjusted to minimize overflows and maximize dynamic range.
result 16-bit Flexpoint tensors closely match 32-bit floating point in training deep networks without tuning.

Asymptotic results for weighted floating bodies are established and used to obtain new proofs for the existence of floating areas on the sphere and in hyperbolic space and to establish the existence of floating areas in Hilbert geometries. Results on weighted best and random approximation and the new approach to floati…

2016-11-13abs ↗pdf ↗

G-Net constructs binary neural networks with high accuracy using randomized binary embeddings.

problem Creating high-accuracy binary neural networks with theoretical guarantees.
method Proposes a novel floating-point G-Net family with randomized binary embeddings and theoretical accuracy guarantees.
result Empirically, G-Net achieves almost 30% higher accuracy on CIFAR-10 compared to prior HDC models.

We carry out a systematic investigation on floating bodies in real space forms. A new unifying approach not only allows us to treat the important classical case of Euclidean space as well as the recent extension to the Euclidean unit sphere, but also the new extension of floating bodies to hyperbolic space. Our main re…

2016-06-24abs ↗pdf ↗

Researchers show NN-based communication algorithms can be implemented on hardware without significant performance loss.

problem Reducing complexity and improving performance of NN-based communication algorithms for practical hardware implementation.
method Implementation of NN-based algorithms in fixed-point arithmetic with quantized weights on specialized hardware (FPGAs, ASICs).
result It is possible to implement NN-based algorithms in fixed-point arithmetic with quantized weights on hardware without significant performance loss.

The study analyzes convergence of adaptive optimizers under low-precision training.

problem Understanding why low-precision training remains effective for large models.
method Developed a theoretical framework for analyzing convergence of adaptive optimizers under floating-point quantization.
result Adaptive optimizers retain convergence rates close to full-precision methods under logarithmic mantissa scaling.

We establish a connection between capillary floating in neutral equilibrium and the billiard ball problem. This allows us to reduce the question of floating in neutral equilibrium at any orientation with a prescribed contact angle for infinite homogeneous cylinders to a question about billiard caustics for their orthog…

2010-12-11abs ↗pdf ↗

Paper examines floating exercise boundaries for American options in time-inhomogeneous models.

problem Floating exercise boundaries in time-inhomogeneous models with negative interest rates or yields.
method Semi-analytical approach for pricing American options.
result Specialized pricing methodologies are required for models with floating exercise boundaries.

Efficient Winograd convolution for INT8 networks using RNS.

problem Difficulty in applying Winograd algorithm to low-precision quantized networks.
method Extends Winograd algorithm to Residue Number System (RNS) for efficient INT8 convolution.
result Arithmetic complexity reduction up to 7.03x with performance improvement up to 2.30x-4.69x.

Deep neural networks struggle with numerical instability during training.

problem Numerical instability in gradient descent training of deep neural networks.
method Analysis of floating-point arithmetic and gradient descent in ReLU neural networks.
result It is highly unlikely for ReLU networks to maintain a superlinear number of affine pieces during training.

This work adds FLOPs to neural network training objectives.

problem Training resource-efficient neural networks with varying FLOPs requirements.
method Extends a state-of-the-art technique to incorporate FLOPs as an optimization objective.
result Different neural networks can be successfully trained for image classification given a desired FLOPs requirement.