Bio-inspired neuromorphic hardware is a research direction to approach brain's computational power and energy efficiency. Spiking neural networks (SNN) encode information as sparsely distributed spike trains and employ spike-timing-dependent plasticity (STDP) mechanism for learning. Existing hardware implementations of…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New KWS neural networks improve accuracy and power efficiency.
Deep learning architectures (DLA) have shown impressive performance in computer vision, natural language processing and so on. Many DLA make use of cloud computing to achieve classification due to the high computation and memory requirements. Privacy and latency concerns resulting from cloud computing has inspired the …
Benchmark for DL inference on embedded HWAs, focusing on autonomous driving.
New Ising models improve consensus clustering on specialized hardware.
DANCE optimizes neural network and accelerator design for faster, more efficient DNN execution.
Recent breakthroughs in Deep Learning (DL) applications have made DL models a key component in almost every modern computing system. The increased popularity of DL applications deployed on a wide-spectrum of platforms have resulted in a plethora of design challenges related to the constraints introduced by the hardware…
Method predicts hardware resource usage by control software with guaranteed linear convergence.
Quantum algorithms for CVaR portfolio optimization face trade-offs between hardware coherence and expressibility.
This paper tackles co-design of neural hardware and software to improve efficiency.
We derive scaling laws for optimizing neural networks in hardware.
New method bounds hardware noise without assumptions.
This paper highlights new opportunities for designing large-scale machine learning systems as a consequence of blurring traditional boundaries that have allowed algorithm designers and application-level practitioners to stay -- for the most part -- oblivious to the details of the underlying hardware-level implementatio…
Applying deep neural networks (DNNs) in mobile and safety-critical systems, such as autonomous vehicles, demands a reliable and efficient execution on hardware. Optimized dedicated hardware accelerators are being developed to achieve this. However, the design of efficient and reliable hardware has become increasingly d…
Optimized neural networks for Edge TPU achieve high accuracy in real-time image classification.
With the rising popularity of machine learning and the ever increasing demand for computational power, there is a growing need for hardware optimized implementations of neural networks and other machine learning models. As the technology evolves, it is also plausible that machine learning or artificial intelligence wil…
VegasFlow accelerates complex simulations across various hardware platforms.
The ever increasing computational cost of Deep Neural Networks (DNN) and the demand for energy efficient hardware for DNN acceleration has made accuracy and hardware cost co-optimization for DNNs tremendously important, especially for edge devices. Owing to the large parameter space and cost of evaluating each paramete…
Paper presents efficient algorithms for convolutional neural networks using Winograd minimal filtering.
Hardware-accelerated RBM solves large combinatorial problems and integer factorization.
Quantum algorithm reduces CVA risk-neutral expectation estimation costs.
Improves VQAs by balancing classical and quantum training resources.
Pipelined Backpropagation trains large models without batches efficiently.
End-to-end performance estimation and measurement of deep neural network (DNN) systems become more important with increasing complexity of DNN systems consisting of hardware and software components. The methodology proposed in this paper aims at a reduced turn-around time for evaluating different design choices of hard…
HotNAS reduces AI search time from hundreds of GPU hours to less than 3 GPU hours.
TensorFlow Probability MCMC toolkit improves MCMC efficiency for modern hardware.
HarDNN detects and protects CNNs from hardware errors.
On-device CNN inference for real-time computer vision applications can result in computational demands that far exceed the energy budgets of mobile devices. This paper proposes FixyNN, a co-designed hardware accelerator platform which splits a CNN model into two parts: a set of layers that are fixed in the hardware pla…
Novel NAS method balances performance and hardware metrics efficiently.
Galen algorithm compresses neural networks for specific hardware with reduced latency.
Large-scale deep neural networks are both memory intensive and computation-intensive, thereby posing stringent requirements on the computing platforms. Hardware accelerations of deep neural networks have been extensively investigated in both industry and academia. Specific forms of binary neural networks (BNNs) and sto…
ALF reduces network parameters and operations by 70% and 61%, respectively, on embedded hardware.
Recurrent neural networks (RNNs) have shown excellent performance in processing sequence data. However, they are both complex and memory intensive due to their recursive nature. These limitations make RNNs difficult to embed on mobile devices requiring real-time processes with limited hardware resources. To address the…
New method uses surrogate gradients to train efficient spiking networks on neuromorphic hardware.
Efficient deep learning computing requires algorithm and hardware co-design to enable specialization: we usually need to change the algorithm to reduce memory footprint and improve energy efficiency. However, the extra degree of freedom from the algorithm makes the design space much larger: it's not only about designin…
Major advancements in building general-purpose and customized hardware have been one of the key enablers of versatility and pervasiveness of machine learning models such as deep neural networks. To sustain this ubiquitous deployment of machine learning models and cope with their computational and storage complexity, se…
DJPQ optimizes neural network pruning and quantization for hardware efficiency.
CoCoPIE shows AI can run on regular devices without special hardware.
NASCaps automates CapsNet design for better accuracy and hardware efficiency.
Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural networks (DNNs). An algorithm-hardware co-optimization framework is developed, which i…
New activation networks improve model efficiency and performance.
Improved neural network training with ADMM for hardware compatibility.
Bayesian Neural Networks (BNNs) have been proposed to address the problem of model uncertainty in training and inference. By introducing weights associated with conditioned probability distributions, BNNs are capable of resolving the overfitting issue commonly seen in conventional neural networks and allow for small-da…
In recent years, Convolutional Neural Network (CNN) based methods have achieved great success in a large number of applications and have been among the most powerful and widely used techniques in computer vision. However, CNN-based methods are computational-intensive and resource-consuming, and thus are hard to be inte…
Improved machine translation with INT8 hardware using a novel training method.
Convolutional neural networks (CNNs) demand huge DRAM bandwidth for computational imaging tasks, and block-based processing has recently been applied to greatly reduce the bandwidth. However, the induced additional computation for feature recomputing or the large SRAM for feature reusing will degrade the performance or…
Predicting MRI coil failures using time series classification.
Game theory shows miners' hardware improvements don't centralize mining.