Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2515027521,003 · Jun 202019922001200920182026
48 results for neural network binarization

Channel pruning and weight binarization improve keyword spotting accuracy.

problem Improving accuracy of keyword spotting in neural networks.
method Group-wise splitting method using group Lasso penalty for channel sparsity, combined with 1-bit weight precision.
result Achieved over 50% channel sparsity with minimal accuracy loss.

Binarized CNNs improve GPU inference efficiency on resource-constrained devices.

problem Efficient inference on resource-constrained devices for image classification.
method Binarization of weights and computations in CNNs, implemented on GPUs.
result 7.4X speedup with 4.4% accuracy loss on embedded GPU platforms.

TentacleNet improves binarized CNNs, reducing accuracy loss and memory usage.

problem Excessive accuracy loss in binarized CNNs.
method Parallelization inspired by ensemble learning theory, end-to-end trainable compact topology.
result Significant memory savings compared to state-of-the-art binary ensemble methods.

This work tackles catastrophic forgetting in neural networks by mimicking brain's metaplasticity.

problem Catastrophic forgetting in neural networks, where new tasks erase previously learned ones.
method Interpreting binarized neural networks as metaplastic systems, adjusting their training technique.
result Training technique reduces catastrophic forgetting without needing previously presented data.

Binary representation is desirable for its memory efficiency, computation speed and robustness. In this paper, we propose adjustable bounded rectifiers to learn binary representations for deep neural networks. While hard constraining representations across layers to be binary makes training unreasonably difficult, we s…

2015-11-19abs ↗pdf ↗

Compact binary CNNs improve landmark localization on limited resources.

problem Designing lightweight architectures for landmark localization with limited resources.
method Binarization of neural networks, hierarchical, parallel, multi-scale residual architecture.
result Significant performance improvement over standard architectures.

Study on function sensitivity in random DNNs using large deviation theory.

problem Understanding function sensitivity in finite-size deep neural networks.
method Large deviation theory and path integral analysis applied to random DNNs with ReLU and sign activations.
result Random DNNs with ReLU activations are more robust to parameter perturbations.

High error rates improve neural network performance and reduce power consumption.

problem Training neural networks with high Bit Error Rates (BERs).
method Trained three Binarized Convolutional Neural Network architectures on various datasets with high BERs.
result High BERs do not significantly degrade test accuracy, enabling more efficient hardware.

QubitHD improves HD computing ML efficiency without sacrificing accuracy.

problem Trade-off between energy efficiency and classification accuracy in HD computing-based ML.
method Stochastically binarizes HD-based algorithms while maintaining comparable classification accuracies.
result 65% improvement in energy efficiency and 95% improvement in training time on FPGA.

Progressive stochastic binarization improves deep network inference efficiency.

problem Efficient inference of deep networks with reduced memory and computational resources.
method A progressive stochastic binarization scheme for deep networks that uses small integers and fixed shifts.
result Matches the accuracy of previous binarized approaches and reduces inference costs by up to 33%.

A method for efficient approximate inference on discrete distributions.

problem Applying SVGD to discrete distributions.
method Transforming discrete distributions to piecewise continuous distributions for SVGD application.
result Outperforms traditional algorithms and ensemble methods on discrete graphical models.

This paper introduces GLT for better input data representation in BNN and proposes a compact topology with block pruning.

problem Improving input data representation for Binary Neural Networks (BNN).
method Generic Learned Thermometer (GLT) for encoding, block pruning and Knowledge Distillation for compact topology.
result Significant accuracy gains and lightweight fully-binarized models with limited accuracy degradation.

This work improves DNN weight quantization with ADMM, achieving lossless binarization and reduced search space.

problem Improving DNN model compression and accuracy with low bit quantization.
method Extending ADMM framework for DNN weight quantization with progressive multi-step approach.
result Achieved lossless and fully binarized DNNs with reduced accuracy loss.

LUTNet optimizes FPGA for neural network inference, achieving high efficiency.

problem Redundancy in neural networks leads to inefficient hardware implementations.
method Exploits LUTs' flexibility for efficient neural network inference on FPGAs.
result Significant area savings with comparable accuracy compared to state-of-the-art binarized neural networks.

Paper proposes efficient BNN inference techniques on FPGA.

problem Redundancy in BNN inference leading to high computation and data access costs.
method Analyzed image similarity and BNN kernel weights to exploit redundancy. Proposed two types of fast and energy-efficient architectures.
result 80% reduction in computation and 40% in buffer access, achieving 17% power reduction.

New approach quantizes neural nets with guaranteed convergence to loss-optimal states.

problem Optimizing deep neural net compression by quantizing weights.
method Model compression as constrained optimization framework, alternating learning and quantization.
result Guaranteed convergence to local optimum of loss for quantized nets, achieving high compression rates.

Gaudy images help train deep neural networks with less data.

problem Training deep neural networks with limited real data from visual cortex neurons.
method Used high-contrast binarized natural images (gaudy images) to train DNNs.
result Reduced training data needed for accurate DNN predictions of visual cortex neuron responses.

New model CRS combines transparency and high performance for classification.

problem Need models with transparent structure and high classification performance.
method Concept Rule Sets (CRS) with Multilayer Logical Perceptron (MLLP) and Random Binarization (RB).
result CRS outperforms state-of-the-art approaches and has low complexity.

Improves bit error tolerance in RRAM-based BNNs without overfitting.

problem Bit errors in RRAM-based BNNs reduce accuracy and overfit to training error rates.
method Proposes straight-through gradient approximation and a novel regularizer.
result Improves BNNs' robustness to bit errors without overfitting.

This paper justifies the use of straight-through estimator in training quantized neural nets.

problem Minimizing loss in quantized neural nets with vanishing gradients.
method Introduced straight-through estimator (STE) and proved its effectiveness in two-linear-layer network with binarized ReLU activations.
result Proved that the coarse gradient derived from STE is a descent direction for minimizing population loss.

Quantized neural networks can improve robustness against adversarial attacks.

problem Adversarial attacks on neural networks with low-precision weights and activations.
method Proposed a third benefit of very low-precision neural networks: improved robustness against some adversarial attacks. Focused on weights and activations quantized to ±1, and conducted black-box and white-box experiments.
result Non-scaled binary neural networks can reduce the impact of iterative attacks, but do not artificially mask gradients.

This research bridges binary and spiking neural networks for efficient on-chip AI.

problem Reducing compute requirements in machine learning frameworks.
method Training Spiking Neural Networks in extreme quantization regime and utilizing standard training techniques for conversion.
result Training Spiking Neural Networks in extreme quantization regime achieves near full precision accuracies.

New approach uses Boolean circuits to optimize neural networks.

problem Improving efficiency of neural network implementations on hardware accelerators.
method Formalized neural networks as Boolean circuits, showing binarized networks are functionally complete.
result Binarized neural networks are functionally complete, suggesting new possibilities for neural network accelerators.

LUTNet optimizes FPGA neural network accelerators by leveraging LUTs for inference, achieving significant area savings.

problem Redundancy in deep neural networks and inefficient use of FPGA resources.
method End-to-end hardware-software framework using LUTs to implement any K-input Boolean operation for inference.
result Significant area savings and comparable accuracy compared to state-of-the-art binarized neural networks.

Paper models entropy-based impact of soft errors on neural network inference.

problem Estimating impact of radiation-induced faults on neural network inference.
method Entropy-based statistical models for SEU and MBU across layers.
result Accurate models to evaluate error-resiliency of neural network topologies.

New framework reduces LLM complexity by directly finetuning in Boolean domain.

problem Reducing the complexity of large language models (LLMs) while maintaining performance.
method Proposes a novel framework using multi-kernel Boolean parameters for direct finetuning in the Boolean domain.
result Significantly reduces complexity during both finetuning and inference, outperforming recent techniques.

CNN detects phase transitions in Potts models without prior knowledge.

problem Detecting phase transitions in qq-state Potts models using deep learning.
method Trained a deep CNN on Ising model spin configurations and temperatures, then tested on Potts model images.
result Deep CNN accurately detects phase transitions in Potts models, including high- and low-temperature regions.

DFM binarizes feature embeddings for fast, accurate recommendation.

problem Expensive storage and computational cost due to large feature dimensions.
method DFM binarizes real-valued model parameters into binary codes for efficient storage and computation.
result DFM outperforms state-of-the-art binarized recommendation models and shows competitive performance compared to its real-valued version.

Sparse binary compression reduces communication costs in distributed deep learning.

problem Limited communication bandwidth in distributed deep learning.
method Combines gradient sparsification, binarization, and optimal weight update encoding.
result Reduces upstream communication by more than four orders of magnitude.

MeliusNet improves binary neural networks to match MobileNet-v1 accuracy.

problem Achieving high accuracy with binary neural networks on mobile devices.
method Alternating DenseBlocks and ImprovementBlocks to increase feature capacity and quality.
result MeliusNet matches MobileNet-v1 accuracy on ImageNet, improving binary network performance.

BinaryDuo improves BNNs by coupling binary activations, outperforming state-of-the-art models.

problem Gradient mismatch in BNNs due to binarizing activations.
method Using gradient of smoothed loss function to estimate gradient mismatch, proposing BinaryDuo scheme with coupled ternary activations.
result BinaryDuo outperforms state-of-the-art BNNs on various benchmarks.

New memory allocation scheme improves image generation performance.

problem Improving episodic and semantic memory representation in neural networks.
method Developed a hierarchical latent variable model with differentiable, locally block allocated latent memory.
result Improved conditional likelihood values on various datasets.

CodNN uses error-correcting codes to make neural networks more resilient to noise.

problem Neural networks are sensitive to noise, especially in critical applications.
method Construct robust neural networks by coding data or internal layers with error-correcting codes.
result Parity codes can guarantee robustness for a wide range of neural networks, including binarized networks.