Echo optimizes LSTM RNN training on GPUs by reducing memory footprint.
problem Memory bottleneck in LSTM RNN training on GPUs.
method Compiler-based feature map recomputation to estimate and balance memory and execution time.
result Average 1.89X reduction in GPU memory footprint.
GPU optimization speeds up large-scale classification tasks.
problem Efficiently training large-scale classification models on GPUs.
method Judecious GPU-optimization principles applied to TRON algorithm.
result Significant speedups for logistic regression and SVM classification.
Two GPU memory management approaches reduce deep learning model memory usage.
problem Limited GPU memory in deep learning systems.
method Two orthogonal approaches exploiting iterative training algorithm to optimize memory usage.
result Up to 34.2% reduction in memory usage without communication overhead.
Optimal GCP selects GCs for ACGs, reducing memory usage.
problem Limited GPU memory in deep learning training.
method Optimal algorithms for selecting Gradient Checkpoints (GCs) for arbitrary computation graphs (ACGs).
result Achieves maximal memory cut-offs for ACGs, outperforming existing methods.
New algorithm reduces costs and latency for large language model inference.
problem Optimizing inference costs and latency for large language models with GPU constraints.
method Formulated as an online scheduling problem with endogenous memory growth, introduced fluid model and WAIT algorithms.
result Reduced costs and latency, especially in near-overloaded and overloaded regimes.
New GPU algorithm boosts machine learning with larger datasets.
problem Limited GPU memory restricts training data size.
method Out-of-core GPU gradient boosting algorithm.
result Training larger datasets on GPUs without accuracy loss.
Paper proposes a GPU-based system for training massive deep learning models in ads systems.
problem Training massive deep learning models with terabyte-scale parameters in ads systems.
method Hierarchical GPU parameter server with 3-layer storage (GPU High-Bandwidth Memory, CPU main memory, SSD).
result 4-node hierarchical GPU parameter server trains a model 2X faster than a 150-node in-memory system.
New GPU kernels boost deep learning speed and memory efficiency.
problem Sparse deep learning matrices are not well-suited for existing sparse kernels.
method Identified favorable properties of sparse matrices from deep learning, developed high-performance GPU kernels for sparse matrix operations.
result 27% of single-precision peak performance on Nvidia V100 GPUs achieved with new kernels.
ProxylessNAS directly optimizes neural architectures for large-scale tasks without proxy tasks.
problem Inefficient and costly neural architecture search for large-scale tasks.
method Directly learns neural architectures for large-scale tasks and hardware platforms without proxy tasks.
result Achieves better performance and efficiency than previous methods.
RFX accelerates and compresses Random Forests for large datasets.
problem Memory bottleneck in proximity matrices limits Random Forest analysis.
method QLORA compression, CPU TriBlock storage, GPU batch sizing, 3D MDS visualization.
result Proximity-based Random Forest analysis on larger datasets is feasible.
Asymmetry PRISM outperforms CPU and GPU solvers for institutional rebalancing.
problem Institutional rebalancing with deadline constraints
method Asymmetry PRISM
result Asymmetry PRISM-CPU is 4.5x to 24.1x faster than the fastest completed reference row in the same lane.
cuRegOT accelerates GPU-based entropic OT solving.
problem Slow convergence and high computational cost of optimal transport on GPUs.
method High-performance GPU solver with algorithmic and architectural optimizations.
result Significant speedups over state-of-the-art solvers.
New GPU algorithm speeds up Gaussian Process analysis.
problem Reducing computational complexity for large spatial datasets.
method Implemented three GPU methods for Vecchia Approximation.
result New GPU method outperforms existing methods.
Introduces BMF for efficient matrix factorization of large data.
problem Efficiently factorizing large scale matrices with limited memory.
method Uses block matrix approach and factorization at a block level.
result Demonstrates faster convergence on large matrices.
KineticSim accelerates financial market simulations 3406x over CPU.
problem Simulating financial markets at scale with multi-agent models is bottlenecked by sequential processing and GPU kernel overhead.
method Formalized and implemented a reusable parallel design pattern for iterative multi-agent reductions in thread-block shared memory.
result Achieved a peak throughput of over 54.7 billion agent-events per second, delivering 3406x speedup over CPU.
KineticSim: A lightweight, high-performance execution engine for real-time market simulators
problem Simulating financial markets at scale with multi-agent models
method Reusable parallel design pattern: persistent, state-carrying clearing for iterative multi-agent reductions
result Reduces per-step critical-path depth from Theta(L+A) to Theta(log L + ceil(A/L))
MIOpen optimizes deep learning operators for GPUs, accelerating research.
problem Optimizing deep learning operators for efficient GPU execution.
method Highly optimized implementations, fusion, auto-tuning, bfloat16 support.
result Accelerates time to discovery in deep learning research.
This paper optimizes neural network training by packing multiple models on a single GPU.
problem Efficiently sharing limited training resources among multiple neural network models.
method Proposes a primitive called 'pack' to jointly train multiple models on a single GPU.
result Significant performance improvements for hyperparameter tuning, up to 40% for two models.
A new parallel algorithm speeds up Hawkes process estimation.
problem Slow maximum likelihood estimation for Hawkes processes.
method Parallel prefix scan for sparse transition matrices.
result Massive speedup with O(N/P) complexity. XGBoost accelerates machine learning on GPUs.
problem Training large datasets efficiently on GPUs.
method Multi-GPU gradient boosting with data compression and end-to-end GPU parallelism.
result Processed 115 million instances in 3 minutes.
SaberLDA learns a large number of topics from big text data on GPUs.
problem Handling large number of topics with existing GPU-based LDA systems.
method Sparsity-aware algorithm, novel data layout, warp-based sampling kernel, sparse count matrix updating.
result SaberLDA learns up to 10,000 topics from billions of tokens in a few hours.
Study shows Transformer and Neural GPU are Turing complete without external memory.
problem Exploring computational power of modern neural network architectures.
method Analyzing computational properties of Transformer and Neural GPU.
result Transformer and Neural GPU are Turing complete without external memory.
A new method increases 3D medical image segmentation accuracy and speed.
problem Training large 3D medical images on GPUs with limited memory.
method Data-swapping method to enlarge GPU memory and avoid patching.
result Improved segmentation accuracy and speed for full-size images.
Efficiently trains large models on limited-memory accelerators.
problem Training large-scale machine learning models on limited-memory accelerators.
method Novel algorithmic building block using primal-dual coordinate methods and duality gap information.
result Order-of-magnitude speedup for training generalized linear models on large datasets.
SliceOut speeds up deep learning training without sacrificing accuracy.
problem Frequent model re-training and large model training workloads in deep learning.
method SliceOut uses dropout-inspired scheme to drop contiguous sets of units at random, leveraging GPU memory layout.
result 10-40% speedups and memory reduction with minimal accuracy loss.
Improves memory efficiency for meta-learning with large images.
problem High memory usage in meta-learning for few-shot classification.
method LITE: episodic training scheme that decomposes task gradients and back-propagates only a random subset of images.
result Achieves state-of-the-art accuracy on real-world and challenging benchmarks.
New GPU-based algorithm for fast optimal transport on brain tractograms.
problem Efficiently comparing and transferring labels in brain tractograms.
method Multiscale algorithm using Sinkhorn divergences on GPU.
result Smooth assignments for label transfer in tractograms.
PruneTrain speeds up neural network training by dynamically pruning weights.
problem Efficiently training large neural networks with high compute and memory costs.
method Structured group-lasso regularization and reconfiguration techniques to reduce weights and model size.
result Achieved 39% reduction in end-to-end training time for ResNet50 on ImageNet.
Compact DNNs increase memory footprint and reduce energy efficiency.
problem Designing compact deep neural networks (DNNs) for improved energy efficiency.
method Evaluation of recently proposed compact DNNs on a Tesla P100 GPU.
result Higher number of activations and memory footprint lead to reduced energy efficiency.
DoRA improves adaptation efficiency for large models by factoring norms and fusing kernels.
problem High-rank DoRA is computationally expensive and infeasible on common GPUs.
method Factored norms and fused Triton kernels to reduce memory and speed up computation.
result Fused implementation is up to 2.0x faster for inference and 1.9x faster for gradient computation.
GPU acceleration for tree boosting improves speed and scalability.
problem Scalability and performance issues in GPU-based tree building algorithms.
method Histogram-based algorithm for approximate split finding on GPUs.
result 7-8 times faster training on CPU and 25 times faster on Xeon server.
This paper proposes an alternative to E2E training for deep networks, reducing memory footprint.
problem High GPUs memory footprint in end-to-end training of deep networks.
method Locally supervised learning with information propagation loss to avoid information collapse.
result The proposed method achieves competitive performance with less than 40% memory footprint compared to E2E training.
aweSOM accelerates SOM clustering for large datasets.
problem Scalability issues in existing SOM implementations for large, multidimensional data.
method CPU/GPU-accelerated Self-organizing Maps (SOM) with ensemble stacking.
result 10-100x speed up and improved memory efficiency for large datasets.
GraphGP: Scalable Gaussian Processes with Vecchia's Approximation
problem Naive Gaussian Process computation limits practical use
method GPU algorithm for Vecchia's approximation
result Linear time and memory requirements for nearly a billion parameters
Extends GENO framework for GPU optimization of constrained ML problems.
problem Constrained optimization in classical machine learning.
method Extends GENO framework to GPU optimization, specifying problems in a modeling language.
result Solvers on GPU outperform state-of-the-art approaches by several orders of magnitude.
ZeRO optimizes memory for training large models, scaling to trillions of parameters.
problem Training models with billions to trillions of parameters is challenging due to limited device memory.
method ZeRO eliminates memory redundancies in data- and model-parallel training, scaling model size proportional to the number of devices.
result ZeRO trains models of up to 13B parameters without model parallelism, achieving super-linear speedup and throughput of 15 Petaflops.
Cyclic Data Parallelism reduces memory usage and balances gradient communications.
problem Training large deep learning models requires efficient parallelism to scale.
method Cyclic Data Parallelism shifts micro-batches from simultaneous to sequential execution, balancing memory and gradient communications.
result Cyclic Data Parallelism reduces total memory usage and balances gradient communications.
NetFuse merges different DNN models with varying weights for faster inference.
problem Inference speed of DNN models with different weights cannot be improved using existing techniques.
method NetFuse merges models with the same architecture but different weights and inputs, replacing operations with more general ones.
result NetFuse can speed up DNN inference time up to 3.6x on a NVIDIA V100 GPU.
Auto-Keras efficiently searches neural architectures with less computation.
problem Expensive computational cost in existing NAS algorithms.
method Bayesian optimization guided by network morphism.
result Framework outperforms state-of-the-art methods on real-world datasets.
A GPU framework speeds up BnB for discrete optimization problems.
problem Optimizing large-scale discrete problems with GPU limitations.
method Parallel BnB nodes in GPU batches, using padding and custom kernels.
result One to two orders of magnitude speedup and zero optimality gap.
Proposes RBGP framework for efficient block sparse neural networks.
problem Efficiently exploit structured sparsity patterns for sparse neural networks on GPU.
method Uses Ramanujan Bipartite Graph Product to generate structured multi-level block sparse neural networks.
result Achieves 5-9x and 2-5x runtime gains over unstructured and block sparsity patterns respectively, while maintaining accuracy.
ORIGAMI accelerates ML algorithms by splitting compute tasks between in-memory and off-chip accelerators.
problem Memory bandwidth bottleneck in ML processing.
method Heterogeneous in-memory accelerators and off-chip compute platform, pattern-matching for compute patterns, computation-splitting compiler.
result ORIGAMI outperforms state-of-the-art accelerators in performance and energy-efficiency.
Binarized CNNs improve GPU inference efficiency on resource-constrained devices.
problem Efficient inference on resource-constrained devices for image classification.
method Binarization of weights and computations in CNNs, implemented on GPUs.
result 7.4X speedup with 4.4% accuracy loss on embedded GPU platforms.
Optimized parallel RNN training reaches up to 845x speedup.
problem Expensive RNN training through back-propagation through time (BPTT).
method Optimized parallel algorithm \opt based on ELM, leveraging GPU shared memory and QR factorization.
result Up to 845x speedup over sequential training and 20x less time to train.
Enhances neural architecture search efficiency and prevents performance collapse.
problem Improving memory efficiency and preventing performance collapse in neural architecture search.
method Employing continuous relaxation strategy and gradient-based optimization for over-parameterized BCNN construction, introducing Confident Learning Rate and partial channel connections.
result NAS-v2 delivers state-of-the-art search efficiency on CIFAR-10 and ImageNet.
TFLMS optimizes neural network training with larger models by rewriting graphs.
problem Training large neural networks with limited accelerator memory.
method Rewriting the computational graph to swap out and in intermediate results to CPU memory.
result Trained ResNet-50 and 3DUnet with significantly larger batch sizes.
Paper explores reducing precision in SVM for faster text classification.
problem Efficiency in multi-class text classification training.
method Comparison of SVM trained with reduced precision (16-bit, half) vs original.
result Reduced precision training maintains text classification accuracy.
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
problem High memory and computational costs of orthogonal momentum updates.
method AuON uses normalized nonlinear scaling and a 'emergency brake' to handle exploding attention logits.
result AuON achieves strong performance without approximate orthogonal matrices, preserving structural alignment and reconditioning.