Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

3469103137 · Jun 202019922001200920182026
48 results for GPU clusters

Poseidon optimizes deep learning training on GPU clusters by reducing network communication.

problem Substantial parameter synchronization over the network in distributed DL implementations.
method Overlap communication and computation, use a hybrid communication scheme.
result Achieves significant speed-ups in DL training on GPU clusters.

New framework quantifies uncertainty in flexible density-based clustering.

problem Uncertainty quantification in clustering with non-parametric density estimation.
method Martingale posterior distributions and density-based clustering.
result Efficient GPU-compatible inference on clustering structures with uncertainty.

Paper proposes a GPU-based system for training massive deep learning models in ads systems.

problem Training massive deep learning models with terabyte-scale parameters in ads systems.
method Hierarchical GPU parameter server with 3-layer storage (GPU High-Bandwidth Memory, CPU main memory, SSD).
result 4-node hierarchical GPU parameter server trains a model 2X faster than a 150-node in-memory system.

aweSOM accelerates SOM clustering for large datasets.

problem Scalability issues in existing SOM implementations for large, multidimensional data.
method CPU/GPU-accelerated Self-organizing Maps (SOM) with ensemble stacking.
result 10-100x speed up and improved memory efficiency for large datasets.

Our system trains ImageNet in 6.6 minutes using 2048 Tesla P40 GPUs.

problem Training large-scale deep neural networks efficiently and accurately.
method Mixed-precision training, extremely large mini-batch size optimization, and optimized all-reduce algorithms.
result Trains ResNet-50 to 75.8% top-1 test accuracy in 6.6 minutes.

Synaptic cluster-driven evolution improves deep neural networks by reducing synapses and clusters.

problem Efficiently synthesizing deep neural networks with fewer synapses and clusters.
method Synaptic cluster-driven genetic encoding scheme.
result Significantly smaller number of synapses and clusters in offspring networks.

NetFuse merges different DNN models with varying weights for faster inference.

problem Inference speed of DNN models with different weights cannot be improved using existing techniques.
method NetFuse merges models with the same architecture but different weights and inputs, replacing operations with more general ones.
result NetFuse can speed up DNN inference time up to 3.6x on a NVIDIA V100 GPU.

End-to-end speech recognition system trained on GPUs and CPUs.

problem Building state-of-the-art speech recognition systems.
method Utilizes CPUs and GPUs for training, data augmentation, and neural network updates. Uses vocal tract length perturbation and acoustic simulator for data augmentation. Employed Horovod allreduce for training.
result Achieved 7.92% WER on proprietary English Bixby open domain test set using a Bidirectional Full Attention (BFA) model.

DSSP improves deep learning training speed by dynamically adjusting staleness thresholds.

problem Time-consuming deep learning training on large datasets.
method Dynamic Stale Synchronous Parallel (DSSP) framework that adapts staleness threshold at runtime.
result DSSP converges faster and achieves higher accuracy than other paradigms.

Stage-based hyper-parameter optimization reduces GPU-hours and training time.

problem Efficiently executing hyper-parameter optimization for deep learning models.
method Stage-based execution strategy to remove redundant computations.
result Stage-based execution outperforms trial-based method by up to 6.60 times in GPU-hours and 4.13 times in training time.

SpinSVAR estimates SVAR models with sparse input, improving accuracy and scalability.

problem Estimating SVAR models with sparse input assumptions.
method SpinSVAR models input as independent Laplacian variables, enforcing sparsity and using least absolute error regression.
result SpinSVAR outperforms state-of-the-art methods in accuracy and runtime, identifying significant structural shocks.

Massively parallel architectures such as the GPU are becoming increasingly important due to the recent proliferation of data. In this paper, we propose a key class of hybrid parallel graphlet algorithms that leverages multiple CPUs and GPUs simultaneously for computing k-vertex induced subgraph statistics (called graph…

2016-08-18abs ↗pdf ↗

New framework for efficient PD averaging and clustering.

problem Challenges in averaging and clustering persistence diagrams.
method Reformulate PD metrics as optimal transport problems, leveraging recent computational advances.
result Scalable computations of PD barycenters and clustering on thousands of diagrams.

EBIC is a biclustering tool for big genomic data, achieving significant speedup.

problem Mining genetic data for high-dimensional and big data challenges.
method EBIC is a biclustering algorithm enhanced for big data, including support for missing values and integration with R.
result EBIC achieves over 6.6 fold speedup on large datasets, demonstrating high scalability.

A GPU framework speeds up BnB for discrete optimization problems.

problem Optimizing large-scale discrete problems with GPU limitations.
method Parallel BnB nodes in GPU batches, using padding and custom kernels.
result One to two orders of magnitude speedup and zero optimality gap.

New GPU kernels boost deep learning speed and memory efficiency.

problem Sparse deep learning matrices are not well-suited for existing sparse kernels.
method Identified favorable properties of sparse matrices from deep learning, developed high-performance GPU kernels for sparse matrix operations.
result 27% of single-precision peak performance on Nvidia V100 GPUs achieved with new kernels.