CoCoPIE shows AI can run on regular devices without special hardware.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
HotNAS reduces AI search time from hundreds of GPU hours to less than 3 GPU hours.
Analog deep learning shows promise but faces scalability challenges.
Article evaluates AI security threats and proposes multiple measures.
Quantum circuits explained using Shapley values for better understanding.
Survey of determinism issues in financial AI systems.
AI progress measured by reduced compute needed to reach past performance.
AI applications pose increasing demands on performance, so it is not surprising that the era of client-side distributed software is becoming important. On top of many AI applications already using mobile hardware, and even browsers for computationally demanding AI applications, we are already witnessing the emergence o…
The MIT/IEEE/Amazon GraphChallenge.org encourages community approaches to developing new solutions for analyzing graphs and sparse data. Sparse AI analytics present unique scalability difficulties. The proposed Sparse Deep Neural Network (DNN) Challenge draws upon prior challenges from machine learning, high performanc…
Despite significant advances in artificial intelligence (AI) for computer vision, its application in medical imaging has been limited by the burden and limits of expert-generated labels. We used images from optical coherence tomography angiography (OCTA), a relatively new imaging modality that measures perfusion of the…
Paper proposes efficient BNN inference flow to reduce computation and memory costs.
FP6 quantization outperforms INT4 in diverse generative tasks for LLMs.
Developed causal chambers for AI validation, providing real-world data.
In order to port the performance of trained artificial neural networks (ANNs) to spiking neural networks (SNNs), which can be implemented in neuromorphic hardware with a drastically reduced energy consumption, an efficient ANN to SNN conversion is needed. Previous conversion schemes focused on the representation of the…
Bio-inspired neuromorphic hardware is a research direction to approach brain's computational power and energy efficiency. Spiking neural networks (SNN) encode information as sparsely distributed spike trains and employ spike-timing-dependent plasticity (STDP) mechanism for learning. Existing hardware implementations of…
New KWS neural networks improve accuracy and power efficiency.
Deep learning architectures (DLA) have shown impressive performance in computer vision, natural language processing and so on. Many DLA make use of cloud computing to achieve classification due to the high computation and memory requirements. Privacy and latency concerns resulting from cloud computing has inspired the …
Benchmark for DL inference on embedded HWAs, focusing on autonomous driving.
Neural network compression methods have enabled deploying large models on emerging edge devices with little cost, by adapting already-trained models to the constraints of these devices. The rapid development of AI-capable edge devices with limited computation and storage requires streamlined methodologies that can effi…
New Ising models improve consensus clustering on specialized hardware.
QTAML models quantum tunneling errors for AI robustness.
DANCE optimizes neural network and accelerator design for faster, more efficient DNN execution.
NEST optimizes deep learning training by placing devices efficiently across networks and memory.
Recent breakthroughs in Deep Learning (DL) applications have made DL models a key component in almost every modern computing system. The increased popularity of DL applications deployed on a wide-spectrum of platforms have resulted in a plethora of design challenges related to the constraints introduced by the hardware…
Method predicts hardware resource usage by control software with guaranteed linear convergence.
Quantum algorithms for CVaR portfolio optimization face trade-offs between hardware coherence and expressibility.
The popularity of Deep Learning for real-world applications is ever-growing. With the introduction of high performance hardware, applications are no longer limited to image recognition. With the introduction of more complex problems comes more and more complex solutions, and the increasing need for explainable AI. Deep…
This paper tackles co-design of neural hardware and software to improve efficiency.
APQ jointly optimizes neural architecture, pruning, and quantization for efficient inference.
We derive scaling laws for optimizing neural networks in hardware.
New method bounds hardware noise without assumptions.
This paper highlights new opportunities for designing large-scale machine learning systems as a consequence of blurring traditional boundaries that have allowed algorithm designers and application-level practitioners to stay -- for the most part -- oblivious to the details of the underlying hardware-level implementatio…
Applying deep neural networks (DNNs) in mobile and safety-critical systems, such as autonomous vehicles, demands a reliable and efficient execution on hardware. Optimized dedicated hardware accelerators are being developed to achieve this. However, the design of efficient and reliable hardware has become increasingly d…
Optimized neural networks for Edge TPU achieve high accuracy in real-time image classification.
With the rising popularity of machine learning and the ever increasing demand for computational power, there is a growing need for hardware optimized implementations of neural networks and other machine learning models. As the technology evolves, it is also plausible that machine learning or artificial intelligence wil…
VegasFlow accelerates complex simulations across various hardware platforms.
The ever increasing computational cost of Deep Neural Networks (DNN) and the demand for energy efficient hardware for DNN acceleration has made accuracy and hardware cost co-optimization for DNNs tremendously important, especially for edge devices. Owing to the large parameter space and cost of evaluating each paramete…
SmallML predicts customer churn for SMEs with small data, improving accuracy by 24.2 points.
Paper presents efficient algorithms for convolutional neural networks using Winograd minimal filtering.
Hardware-accelerated RBM solves large combinatorial problems and integer factorization.
Quantum algorithm reduces CVA risk-neutral expectation estimation costs.
Improves VQAs by balancing classical and quantum training resources.
Pipelined Backpropagation trains large models without batches efficiently.
Resource-constrained IoT devices, such as sensors and actuators, have become ubiquitous in recent years. This has led to the generation of large quantities of data in real-time, which is an appealing target for AI systems. However, deploying machine learning models on such end-devices is nearly impossible. A typical so…
End-to-end performance estimation and measurement of deep neural network (DNN) systems become more important with increasing complexity of DNN systems consisting of hardware and software components. The methodology proposed in this paper aims at a reduced turn-around time for evaluating different design choices of hard…
In recent years, advances in deep learning have resulted in unprecedented leaps in diverse tasks spanning from speech and object recognition to context awareness and health monitoring. As a result, an increasing number of AI-enabled applications are being developed targeting ubiquitous and mobile devices. While deep ne…
We use GG distributions to optimize LLMs, reducing size and training time.
HarDNN detects and protects CNNs from hardware errors.