Paper reduces AI complexity with pre-defined sparsity and hardware acceleration.
problem Reduction of computational and storage complexity in neural networks.
method Pre-defined sparsity and hardware acceleration architecture.
result Significant reduction in storage and computational complexity (5X+ reduction) without significant performance loss.
Researchers show NN-based communication algorithms can be implemented on hardware without significant performance loss.
problem Reducing complexity and improving performance of NN-based communication algorithms for practical hardware implementation.
method Implementation of NN-based algorithms in fixed-point arithmetic with quantized weights on specialized hardware (FPGAs, ASICs).
result It is possible to implement NN-based algorithms in fixed-point arithmetic with quantized weights on hardware without significant performance loss.
Graph neural networks improve charged particle tracking on FPGAs.
problem Charged particle trajectory determination in high interaction density conditions.
method Graph neural networks (GNNs) embedded in tracker data as graphs, classifying edges as track segments.
result GNNs implemented on FPGAs for charged particle tracking, enabling future HL-LHC experiments.
A hardware-based reservoir computing system predicts time series with high speed and accuracy.
problem Processing time-dependent signals with high speed and accuracy.
method A hardware-based reservoir computing system using a field-programmable gate array (FPGA) for both the reservoir and output layers.
result Achieves comparable accuracy to software approaches but with a superior real-time prediction rate up to 160 MHz.
New online adaptive SVR model for IVS with hardware acceleration.
problem Modeling implied volatility surface (IVS) in real-time.
method Online adaptive primal support vector regression (SVR) with hardware acceleration.
result Gaussian kernel outperforms linear kernel in support vector size regulation.
Paper presents FPGA implementation for efficient recurrent neural networks.
problem Implementing recurrent neural networks on FPGAs for low latency.
method Developed hls4ml framework to implement LSTM and GRU layers.
result Demonstrated effective designs for both small and large models.
Accelerator boosts energy efficiency for MANNs on FPGAs.
problem Efficiently running MANNs on accelerators designed for other NNs.
method Data flow architecture, inference thresholding.
result Higher energy efficiency compared to NVIDIA GPU.
Real-time semantic segmentation for autonomous vehicles on FPGA reduces latency and power consumption.
problem Efficient real-time semantic segmentation for autonomous vehicles.
method Compressed ENet architecture, FPGA deployment, batch processing, filter reduction, quantization-aware training.
result Reduced latency to 3 ms per image with batch size of ten and 40% resource utilization.
New framework reduces cost of financial option pricing simulations on FPGAs.
problem Efficiently simulate financial option pricing with reduced computational cost.
method Nested MLMC framework with low precision calculations on FPGAs.
result Higher computational savings compared to existing mixed-precision MLMC frameworks.
LeFlow enables quick FPGA synthesis from Tensorflow models.
problem Manual translation of Tensorflow models to FPGA RTL is time-consuming and requires expertise.
method Uses XLA compiler to emit synthesizable LLVM code from Tensorflow specifications, which is then synthesized.
result Allows users to generate Deep Neural Networks with just a few lines of Python code.
Kolmogorov-Arnold Networks enable ultrafast online learning with fixed-point quantization.
problem Efficient online learning for high-frequency systems with strict memory constraints.
method Fixed-point online training on FPGAs exploiting B-spline locality in KANs.
result Kolmogorov-Arnold Networks are more efficient and expressive than MLPs for low-latency tasks.
Efficient FPGA dropout algorithm reduces memory usage.
problem Overfitting in deep neural networks.
method Hardware-oriented dropout algorithm for FPGA implementation.
result Significant resource reduction in FPGA implementation.
PolyLUT uses polynomials to reduce FPGA latency.
problem Reducing latency in FPGA-based neural network inference.
method Training neural networks using multivariate polynomials as basic building blocks.
result Achieved significant latency and area improvements.
New approach uses Boolean circuits to optimize neural networks.
problem Improving efficiency of neural network implementations on hardware accelerators.
method Formalized neural networks as Boolean circuits, showing binarized networks are functionally complete.
result Binarized neural networks are functionally complete, suggesting new possibilities for neural network accelerators.
PoET-BiN reduces power consumption in neural networks on embedded devices.
problem Power inefficiency in neural network implementations on embedded platforms.
method Look-Up Table based implementation with a modified Decision Tree approach.
result Near state-of-the-art results with up to 6 orders of magnitude energy reduction.
A new method improves RVFL networks for resource-efficient machine learning.
problem Deploying machine learning on edge devices with limited resources.
method Density encoding and hyperdimensional computing operations.
result The proposed method achieves higher accuracy and lower energy consumption.
Trading system uses NP-hard optimization to select stocks for high Sharpe ratio trading.
problem Finding profitable, uncorrelated stocks for high Sharpe ratio trading.
method NP-hard combinatorial optimization using Ising machine and simulated bifurcation algorithm.
result Trading strategy with FPGA-based system achieves 164 μs response latency.
This paper optimizes deep learning systems for high performance and low energy consumption.
problem Achieving ultra-high energy efficiency and performance for deep neural networks.
method Developed an algorithm-hardware co-optimization framework that reduces computational and storage complexity.
result Achieved at least 152X speedup and 71X energy efficiency gain compared to IBM TrueNorth processor.
FPGAs accelerate DNNs with binarized weights, improving power and speed.
problem Power and speed limitations of GPU-accelerated DNNs.
method Binarized neural networks on FPGAs using OpenCL.
result Near state-of-the-art performance with >16x power savings.
New methodology optimizes complex systems design trade-offs.
problem Complex multi-objective optimization with unknown feasibility constraints.
method Introduces HyperMapper 2.0 for multi-objective optimization, handling categorical/ordinal variables.
result HyperMapper 2.0 provides better Pareto fronts and 8x improvement in sampling budget.
Paper proposes NASAIC framework for co-designing neural architectures and heterogeneous ASICs.
problem Designing efficient neural architectures and ASICs for multiple tasks.
method Build ASIC templates and propose NASAIC framework for simultaneous design of architectures and ASICs.
result NASAIC ensures design specifications and maximizes accuracy with minimal performance loss.
The rapid growth of data size and accessibility in recent years has instigated a shift of philosophy in algorithm design for artificial intelligence. Instead of engineering algorithms by hand, the ability to learn composable systems automatically from massive amounts of data has led to ground-breaking performance in im…
Improves magnetic field mapping using an array of magnetometers with noisy input.
problem Improving magnetic field maps in indoor environments with noisy magnetometer data.
method Uses Gaussian process regression with an array of magnetometers, incorporating known array positions and relative magnetometer locations.
result The method produces higher quality magnetic field maps compared to using a single magnetometer.
NeuraLUT maps neural networks to lookup tables, reducing latency and improving expressivity.
problem Reducing latency in deep neural networks for FPGA accelerators.
method Mapping entire sub-networks to a single lookup table, introducing skip connections.
result Up to 4.3x lower latency for the same accuracy.
SNRA combines power-efficient probabilistic and deterministic computing for deep belief networks.
problem Efficiently training and evaluating deep belief networks with low power consumption.
method Developed a spintronic neuromorphic reconfigurable array (SNRA) for in-circuit training and evaluation of deep belief networks (DBNs). Used probabilistic spin logic devices and a four-state finite state machine for unsupervised training.
result SNRA achieves more than 80% reduction in combined dynamic and static power dissipation compared to SRAM-based configurable fabrics.
Channel gating reduces CNN computation cost by skipping ineffective feature regions.
problem Reducing computation cost in CNNs while maintaining accuracy.
method Dynamic, fine-grained pruning scheme that identifies and skips computation on ineffective feature regions.
result 2.7-8.0x reduction in FLOPs and 2.0-4.4x reduction in memory accesses with minimal accuracy loss.
We discuss of the conceptual difficulties connected with the anticommutativity of classical fermion fields, and we argue that the "space" of all classical configurations of a model with such fields should be described as an infinite-dimensional supermanifold M. We discuss the two main approaches to supermanifolds, and …
Combines gating and tensor products for RNNs to improve performance.
problem Improving RNNs' ability to capture long-term dependencies.
method Proposes a novel RNN architecture combining gating mechanism and tensor products.
result Significant performance improvement on word-level and character-level language modeling tasks.
Researchers mapped CNNs to resistive devices for training, overcoming noise and bound limitations.
problem Training deep CNNs with resistive cross-point devices.
method Mapped CNN layers to RPU arrays, implemented noise and bound management techniques, and digitally programmable update management.
result Successfully applied RPU concept for training CNNs, enabling broader applicability.
This paper designs sensor arrays for estimating unsteady flows efficiently.
problem Estimating high-dimensional unsteady flow fields with limited sensor placement.
method Combines data-driven modeling, Kalman Filter design, and sparsification for sensor selection.
result Proposed sensor arrays are highly effective for flow-field estimation across various conditions.
The paper proves existence of solutions for mean field equations on compact Riemann surfaces.
problem Existence of solutions for mean field equations on compact Riemann surfaces.
method Min-max scheme introduced by Djadli-Malchiodi (2006) and Djadli (2008).
result Proves existence of solutions for mean field equations on compact Riemann surfaces.
Gating units in GRUs and LSTMs create slow modes and control phase-space complexity.
problem Training challenges in RNNs due to exploding or vanishing gradients.
method Random matrix theory and mean-field theory applied to GRUs and LSTMs.
result Gates in GRUs and LSTMs lead to accumulation of slow modes and control phase-space complexity.
Paper solves curvature prescription problem on surfaces with boundary.
problem Prescribing Gaussian and geodesic curvatures on compact surfaces with boundary.
method Mean field-type formulation and variational techniques.
result Existence results for positive, zero, and negative Euler characteristics.
Theory explains how recurrent networks remember sequences.
problem Understanding how recurrent networks remember sequences and perform well.
method Mean field theory and random matrix theory applied to RNNs with gating mechanisms.
result Gated RNNs outperform non-gated RNNs in remembering sequences.
Simple construction for universal quantum gates.
problem Designing efficient quantum gates for topological computers.
method Demonstrated a simple construction for unitary solutions of the braided Yang-Baxter equation in any dimension.
result Proved the existence of universal quantum gates in any dimension.
Cognitive radar selects optimal antenna subarrays using deep learning.
problem Optimize radar antenna selection for cost and performance.
method Convolutional Neural Network (CNN) for multi-class classification.
result CNN provides 22% better classification performance and 72% more accurate DoA estimates.
Enhances quantum sensing by eliminating multiple oscillations in field amplitude estimation.
problem Multiple oscillations in field amplitude estimation due to inter-qubit interactions at high qubit densities.
method Adopting a quantum circuit learning framework to approximate a target function by optimizing gate parameters.
result Elimination of multiple oscillations, leading to enhanced dynamic range of quantum sensing.
Tangent automates derivatives in Python, improving expressiveness and performance.
problem Efficiently calculating derivatives for complex models in Python.
method Source-code transformation for dynamically typed array programming.
result Demonstrates improved expressiveness and performance in automatic differentiation.
Study evaluates training programs for unemployed in Belgium using machine learning.
problem Determining which training programs are most effective for unemployed individuals in Belgium.
method Used Modified Causal Forests, a causal machine learning estimator, to analyze data from unemployed in Belgium.
result There is significant heterogeneity in the effectiveness of different training programs for unemployed individuals in Belgium.
A neural network, IHT-Net, improves DOA estimation with sparse arrays.
problem Single-snapshot DOA estimation with sparse arrays in dynamic settings.
method IHT-inspired neural network with recurrent neural network and autoencoders.
result IHT-Net achieves faster convergence and higher accuracy in DOA estimation.
GHNet improves graph learning by balancing homogeneity and heterogeneity.
problem Over-smoothing in GCN leads to similar node representations.
method GHNet uses gating units to balance homogeneity and heterogeneity in feature propagation.
result GHNet achieves larger receptive fields without over-smoothing.
Neural Programmer learns natural language queries for databases.
problem Natural language interface learning for database queries.
method Enhanced Neural Programmer model trained on weak supervision.
result Single Neural Programmer model achieves 34.2% accuracy.
New fault-tolerant quantum gates for homological LDPC codes with constant or almost-constant rate.
problem Fault-tolerant quantum computing for homological LDPC codes with constant or almost-constant encoding rate.
method Derive generic formula for transversal and logical gates acting on 3-manifolds, using higher symmetries and cup product cohomology.
result Parallelizable logical gates for homological LDPC codes with constant or almost-constant rate.
New graph learning model can approximate any function and handle edge values.
problem Graph learning models' limitations in approximating functions and handling edge values.
method Proposes a Graph Neural Network that can approximate any function and handle arbitrary edge values.
result Proves the model is strictly more expressive than existing models.
FixyNN improves energy efficiency of mobile computer vision tasks.
problem High energy consumption of state-of-the-art CNN models on mobile devices.
method Fixed-weight feature extractor and programmable CNN accelerator for transfer learning.
result Achieved up to 26.6 TOPS/W energy efficiency, nearly 2x more efficient than conventional accelerators.
New method forecasts workforce reintegration success rates.
problem Estimating success of reskilling programs in changing labor markets.
method Uses current workforce demand and supply factors, not historical data.
result Average error of 3.9% compared to 5.4% for best benchmark.
Deep neural networks have achieved impressive supervised classification performance in many tasks including image recognition, speech recognition, and sequence to sequence learning. However, this success has not been translated to applications like question answering that may involve complex arithmetic and logic reason…
This paper introduces an efficient method for optimizing deep learning hyperparameters.
problem The high dependency of deep learning algorithms on hyper-parameters.
method Orthogonal Array Tuning Method (OATM) for deep learning hyper-parameter tuning.
result The proposed OATM method significantly saves tuning time compared to state-of-the-art methods.