Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2605207791,039 · Jun 202019922001200920182026
48 results for neural network acceleration

Improved NiNo networks accelerate Adam training by up to 50%.

problem Accelerating neural network training with stable and efficient methods.
method Proposed NiNo networks that leverage neuron connectivity and graph neural networks to nowcast parameters periodically during Adam training.
result Accelerates Adam training by up to 50% in vision and language tasks.

DANCE optimizes neural network and accelerator design for faster, more efficient DNN execution.

problem Challenges in optimizing neural network and accelerator design for efficient DNN execution.
method Differentiable approach to co-exploration of accelerator and network architecture design.
result Significantly shorter time to achieve superior accuracy and hardware cost metrics.

Smartfluidnet accelerates Eulerian fluid simulation with neural networks.

problem Current neural network methods for Eulerian fluid simulation lack flexibility and generalization.
method Smartfluidnet automates model generation and dynamic switching to meet user requirements.
result Smartfluidnet achieves 1.46x and 590x speedup compared to state-of-the-art models, with better simulation quality.

Interpolatron accelerates deep neural network optimization faster than existing methods.

problem Accelerating nonconvex optimization for deep neural networks.
method Proposes Interpolatron, a new interpolation scheme to accelerate nonconvex optimization.
result Interpolatron converges much faster than state-of-the-art methods on DNNs of great depths.

Optimized neural networks for Edge TPU achieve high accuracy in real-time image classification.

problem Designing neural networks for hardware accelerators to achieve optimal performance.
method Hardware-aware neural architecture search and model customization for Edge TPU.
result Improved accuracy-latency tradeoff on Pixel 4's Edge TPU compared to existing models.

Graph neural networks speed up nonnegative matrix factorization.

problem Efficiently factorize nonnegative matrices for various applications.
method Developed a graph neural network that combines bipartite self-attention with ADMM updates.
result Significant acceleration achieved in nonnegative matrix factorization.

New approach uses Boolean circuits to optimize neural networks.

problem Improving efficiency of neural network implementations on hardware accelerators.
method Formalized neural networks as Boolean circuits, showing binarized networks are functionally complete.
result Binarized neural networks are functionally complete, suggesting new possibilities for neural network accelerators.

Improved training of large-scale neural networks with reduced variance noise.

problem Training large-scale neural networks with high variance noise.
method Stochastic variance reduced Nesterov's Accelerated Quasi-Newton method (SVR-NAQ).
result Improved performance compared to conventional methods on benchmark problems.

Improved SDE-BNN model reduces NFEs and accelerates convergence.

problem High computational cost and convergence instability in SDE-BNNs.
method Nesterov's Accelerated Gradient (NAG) method integrated into SDE-BNN framework.
result Significantly reduced number of function evaluations (NFEs) and improved predictive accuracy.

DeepLight accelerates CTR predictions in ad serving by 46X.

problem Significantly increased serving delay and high memory usage for ad serving.
method Explicitly searching feature interactions, pruning layers, promoting sparsity.
result Accelerates model inference by 46X on Criteo dataset.

SPP prunes CNN weights probabilistically for faster inference.

problem Efficiently accelerate Convolutional Neural Networks (CNNs) without significant accuracy loss.
method Structured Probabilistic Pruning (SPP) with adjustable pruning probabilities.
result 4x speedup with minimal accuracy loss (0.3% for AlexNet, 0.8% for VGG-16).

Accelerates deep neural network training with a generalized BN approach.

problem Conventional Batch Normalization (BN) struggles with convergence speed and error rate.
method Introduces Generalized Batch Normalization (GBN) using alternative deviation measures and statistics.
result GBN accelerates training and often improves error rate compared to conventional BN.

Paper reduces AI complexity with pre-defined sparsity and hardware acceleration.

problem Reduction of computational and storage complexity in neural networks.
method Pre-defined sparsity and hardware acceleration architecture.
result Significant reduction in storage and computational complexity (5X+ reduction) without significant performance loss.

Survey of Graph Neural Networks for efficient computation.

problem Efficient processing of Graph Neural Networks (GNNs) is challenging.
method Review of GNN algorithms, software and hardware acceleration analysis.
result Distilled hardware-software, graph-aware, and communication-centric vision for GNN accelerators.

Paper proposes a new binary quantization method for faster DNN inference.

problem Accelerating deep neural network inference on resource-limited devices.
method Quantized Compressed Sensing (QCS) for binary quantization.
result The proposed method preserves benefits of standard methods while reducing quantization error.

Polyak's momentum accelerates training of neural networks.

problem Understanding and explaining the acceleration effect of Polyak's momentum in neural network training.
method Modular analysis of Polyak's momentum for training wide ReLU networks and deep linear networks.
result Polyak's momentum achieves an accelerated linear rate of (1Θ(1κ))t(1-Θ(\frac{1}{\sqrt{κ'}}))^t for training wide ReLU networks and deep linear networks.

AutoAssist accelerates deep neural network training by filtering out less informative instances.

problem Efficiently training deep neural networks with millions of instances.
method AutoAssist uses a shrinking operation to filter out less informative instances, accelerating training.
result AutoAssist reduces training time by 40% for ResNet and 30% for transformer models.

Improves deep neural network training by optimizing activation function and initialization.

problem Inappropriate activation function selection can lead to poor training performance.
method Comprehensive theoretical analysis of the Edge of Chaos and tuning of initialization parameters and activation functions.
result Training acceleration and improved performance achieved by optimizing activation function and initialization.

PruneTrain speeds up neural network training by dynamically pruning weights.

problem Efficiently training large neural networks with high compute and memory costs.
method Structured group-lasso regularization and reconfiguration techniques to reduce weights and model size.
result Achieved 39% reduction in end-to-end training time for ResNet50 on ImageNet.

GRU models with Adam optimizer outperform other combinations in stock market forecasting.

problem Comparing optimization techniques for time series forecasting in LSTM and GRU networks.
method Examined Adam and Nesterov Accelerated Gradient (NAG) on LSTM and GRU models for stock market forecasting.
result GRU models with Adam optimizer produced the lowest RMSE and outperformed other combinations.

New Bayesian model injects noise to improve neural network sparsity and acceleration.

problem Improving neural network sparsity and acceleration.
method Proposes a new Bayesian model that injects noise to neurons outputs while keeping weights unregularized, using log-normal multiplicative noise.
result Provides significant acceleration on deep neural architectures.

VIBNN accelerates Bayesian Neural Networks on FPGAs for efficient inference.

problem Overfitting and small-data training issues in BNNs.
method Hardware accelerator design for variational inference on BNNs, using novel Gaussian random number generators.
result VIBNN achieves high throughput and energy efficiency on FPGA, matching software performance.

Interneurons improve learning in neural networks by accelerating convergence.

problem Rapid adaptation to changing input statistics in neural networks.
method Two mathematically tractable recurrent linear neural networks were compared: one with direct recurrent connections and the other with interneurons that mediate recurrent communication.
result The network with interneurons converges more quickly than the network with direct recurrent connections, scaling logarithmically with initialization spectrum.

Coarse-grained pruning improves sparsity efficiency without sacrificing accuracy.

problem Efficiency of hardware design and prediction accuracy in sparse CNNs.
method Quantitative analysis of sparsity regularity vs. accuracy trade-off.
result Coarse-grained pruning achieves similar sparsity ratios and accuracy as fine-grained pruning.

Paper tackles noisy neural networks and proposes a method to enhance their robustness.

problem Noisy neural networks struggle with random continuous noise in weights.
method Knowledge distillation combined with noise injection during training.
result Models achieve up to twice greater noise tolerance.

GD and NAG accelerate matrix factorization and neural networks.

problem Optimizing rectangular matrix factorization and linear neural networks.
method Gradient descent and Nesterov's accelerated gradient with specific initialization.
result NAG achieves the best-known iteration complexity for these problems.

New insights into GNN optimization reveal skip connections and depth accelerate training.

problem Understanding and optimizing the training of Graph Neural Networks (GNNs).
method Analysis of gradient dynamics in linearized GNNs and empirical validation.
result GNNs are implicitly accelerated by skip connections, more depth, and good label distribution during training.

New DNN method accelerates image processing optimization.

problem Optimizing large-scale inverse problems in image processing.
method Trains a deep neural network to learn parameters for scaled gradient projection method.
result Significantly improves convergence rate of optimization methods.