Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1122 · Nov 201919922001200920182026
16 results for MXNet

A new batch construction method for RNNs outperforms existing approaches in MXNet.

problem Improving the efficiency and performance of recurrent neural networks in MXNet.
method Proposes an alternately sorted batch construction strategy for RNNs.
result Alternately sorted batches outperform bucketing and other methods in training time and recognition performance.

Article presents QR and LQ decomposition algorithms for various matrix sizes and ranks.

problem Solving least squares problems in machine learning and computer vision.
method Developed novel matrix backpropagation algorithms for QR and LQ decompositions of different matrix sizes and ranks.
result Numerical stability and computational efficiency of the proposed methods.

Adds recursion to deep learning frameworks for better handling of recursive data structures.

problem Lack of support for recursion in existing deep learning frameworks.
method Complements existing frameworks with recursive execution of dataflow graphs and APIs for recursive definitions.
result Recursive implementation reduces training and inference time by more effectively using resources.

Sockeye is an open-source toolkit for neural machine translation.

problem Improving Neural Machine Translation (NMT) models and techniques.
method Scalable training and inference for three NMT architectures, including attentional, self-attentional, and fully convolutional networks.
result Sockeye achieves competitive BLEU scores across different NMT architectures, including a best score for its transformer implementation.

AsyB-ProxSGD parallelizes model updates and stochastic gradient descent for large models and data.

problem Efficiently training large models and handling large datasets in parallel.
method AsyB-ProxSGD: model parallel proximal stochastic gradient algorithm for asynchronous systems.
result Achieves linear speedup with O(K1/4)O(K^{1/4}) number of workers for nonconvex problems.

DL2 uses deep learning to optimize resource allocation in DL clusters.

problem Efficient resource scheduling for deep learning clusters is challenging.
method DL2 combines supervised learning and reinforcement learning to dynamically allocate resources.
result DL2 reduces average training completion time by 44.1% compared to fairness scheduler.

Benanza speeds up DL model optimization by automatically generating micro-benchmarks and identifying inefficiencies.

problem Slow characterization/optimization cycles for DL models on GPUs.
method Benanza includes a model processor, benchmark generator, database, and analyzer.
result Benanza identifies optimizations in parallel layer execution, cuDNN, framework inefficiency, layer fusion, and Tensor Cores.

DLBricks automates DL benchmarking on CPUs, reducing effort and time.

problem Lack of representative and up-to-date DL benchmarks on CPUs.
method Decomposes DL models into runnable networks, leveraging layer repetition and auto-generation.
result Accurately estimates DL model performance and speeds up benchmarking time.

Spectral learning extends matrix methods to tensors for better latent variable modeling.

problem Limitations of matrix-based spectral methods in capturing non-Gaussian data.
method Extend spectral decomposition to tensor-based methods for higher-order moments.
result Tensor decomposition can identify latent effects missed by matrix methods.