Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

0.3%0.5%0.8%0.3% · Nov 201819922001200920182026
48 results for Model-Parallelism

When training large machine learning models with many variables or parameters, a single machine is often inadequate since the model may be too large to fit in memory, while training can take a long time even with stochastic updates. A natural recourse is to turn to distributed cluster computing, in order to harness add…

2014-06-18abs ↗pdf ↗

In real world industrial applications of topic modeling, the ability to capture gigantic conceptual space by learning an ultra-high dimensional topical representation, i.e., the so-called "big model", is becoming the next desideratum after enthusiasms on "big data", especially for fine-grained downstream tasks such as …

2014-11-10abs ↗pdf ↗

Mesh-TensorFlow enables efficient deep learning on large clusters.

problem Memory constraints and inefficiency in batch-splitting for large models.
method Introduces Mesh-TensorFlow for specifying general tensor computations across a multi-dimensional mesh of processors.
result Trains Transformer models with up to 5 billion parameters on TPU meshes of up to 512 cores.

ZeRO optimizes memory for training large models, scaling to trillions of parameters.

problem Training models with billions to trillions of parameters is challenging due to limited device memory.
method ZeRO eliminates memory redundancies in data- and model-parallel training, scaling model size proportional to the number of devices.
result ZeRO trains models of up to 13B parameters without model parallelism, achieving super-linear speedup and throughput of 15 Petaflops.

This work optimizes deep learning training by combining data and model parallelism.

problem Training large models with multiple GPUs suffers from high communication overhead and statistical efficiency loss.
method Hybrid parallelization combining data and model parallelism.
result Hybrid training provides significant speedup compared to data parallelism alone.

TSSM splits neural networks for parallel training with minimal accuracy loss.

problem Accuracy degradation in parallel training of deep neural networks.
method TSSM reformulates alternating minimization to achieve parallelism with minimal accuracy loss.
result TSSM achieves significant speedup without accuracy loss on multiple datasets.

Paper proposes DSP to accelerate deep learning model training by addressing locking and straggler problems.

problem Training deep neural networks is slow and inefficient due to locking and straggler issues.
method Layer-wise Staleness and DSP algorithm to handle locking and straggler problems.
result DSP achieves significant training speedup with stronger robustness than compared methods.

Adaptive batch size schedules improve language model training efficiency and generalization.

problem Dilemma of choosing batch sizes in large-scale model training.
method General-purpose adaptive batch size schedules compatible with data and model parallelism.
result Adaptive batch size schedules outperform constant batch sizes and heuristic warmup schedules.

Cyclic Data Parallelism reduces memory usage and balances gradient communications.

problem Training large deep learning models requires efficient parallelism to scale.
method Cyclic Data Parallelism shifts micro-batches from simultaneous to sequential execution, balancing memory and gradient communications.
result Cyclic Data Parallelism reduces total memory usage and balances gradient communications.

The paper optimizes dynamic scheduling for ring architectures in deep learning training.

problem Optimizing deep learning training times with ring architectures.
method Formulated a non-convex, non-linear, NP-hard integer programming problem and developed a doubling heuristic.
result Dynamic scheduling can significantly reduce job completion times in ring architectures.

This paper optimizes how deep learning models are distributed across different devices.

problem Optimizing how large, complex neural networks are split across multiple devices.
method Identified and solved an optimization problem for device placement of DNN operators.
result Automated algorithms that solve the device placement problem for modern pipelined settings.

Paper proposes DCT for efficient hybrid parallel training of large recommendation models.

problem Training large recommendation models at scale with efficient communication.
method Dynamic Communication Thresholding (DCT) for both Data Parallelism and Model Parallelism.
result Reduces communication by 100x and 20x during DP and MP, respectively, improving training time by 37%.

Training examples are not all equally informative. Active learning strategies leverage this observation in order to massively reduce the number of examples that need to be labeled. We leverage the same observation to build a generic strategy for parallelizing learning algorithms. This strategy is effective because the …

2013-10-30abs ↗pdf ↗

IST trains neural networks locally, reducing memory and communication costs.

problem Challenges in distributed learning due to mandatory data separation.
method Independent subnet training (IST) decomposes the network into narrow subnetworks.
result IST reduces training times compared to common distributed learning approaches.

This work proposes an efficient autoregressive model for text generation.

problem The challenge of generating high-quality text with autoregressive models.
method Introduces a cascaded decoding approach using Markov transformers to achieve sub-linear parallel time generation.
result Shows competitive accuracy/speed tradeoff compared to existing methods on five machine translation datasets.

Two parallel samplers enhance image quality in limited denoising steps.

problem Limited denoising steps in diffusion models reduce image quality.
method Two parallel samplers denoise at successive times, integrating their information.
result Two parallel samplers improve image quality compared to a single sampler.

PETRA enables parallel training of deep models with reversible architectures.

problem Challenges in parallelizing deep model training.
method Introduces PETRA, a novel approach for parallelizing gradient computations in reversible architectures.
result Achieves competitive accuracies on CIFAR-10, ImageNet32, and ImageNet using ResNet models.

FasTR efficiently solves sparse and unit-rank tensor regression problems.

problem Sparse and unit-rank tensor regression problems in tensor data analysis.
method FasTR decomposes tensor coefficients into component vectors and estimates each with 1\ell_1 regularized regression, solving in parallel.
result FasTR computes better solutions faster than baseline models.

There is significant recent interest to parallelize deep learning algorithms in order to handle the enormous growth in data and model sizes. While most advances focus on model parallelization and engaging multiple computing agents via using a central parameter server, aspect of data parallelization along with decentral…

2017-06-23abs ↗pdf ↗

Stochastic variational inference (SVI), the state-of-the-art algorithm for scaling variational inference to large-datasets, is inherently serial. Moreover, it requires the parameters to fit in the memory of a single processor; this is problematic when the number of parameters is in billions. In this paper, we propose e…

2016-05-31abs ↗pdf ↗

New method solves blind inverse problems by optimizing both operator and image parameters.

problem Solving blind inverse problems with known forward operator.
method Parallel reverse diffusion guided by gradients from intermediate stages.
result State-of-the-art performance on blind deblurring and imaging through turbulence.

ShadowSync separates background synchronization for scalable distributed training.

problem Reducing synchronization overhead in distributed training for high scalability.
method Separates synchronization from training and runs it in the background.
result Achieves both high throughput and excellent model quality at scale.

What is a systematic way to efficiently apply a wide spectrum of advanced ML programs to industrial scale problems, using Big Models (up to 100s of billions of parameters) on Big Data (up to terabytes or petabytes)? Modern parallelization strategies employ fine-grained operations and scheduling beyond the classic bulk-…

2013-12-30abs ↗pdf ↗

Improves Bayesian neural learning efficiency with surrogate-assisted parallel tempering.

problem Challenges in Bayesian neural learning due to large models and data.
method Combines parallel tempering MCMC with surrogate-assisted optimization for computationally expensive models.
result Significantly lowers computational cost while maintaining quality in decision making.

A new parallel clustering method improves speed and accuracy for single cell transcriptomic data.

problem Challenges in clustering single cell transcriptomic data, including poor quality, lack of prior knowledge, and slow computation.
method Parallel Split Merge Sampling on Dirichlet Process Mixture Model (Para-DPMM).
result The Para-DPMM model outperforms existing methods in clustering quality and computational speed.

Recent technological development has enabled researchers to study social phenomena scientifically in detail and financial markets has particularly attracted physicists since the Brownian motion has played the key role as in physics. In our previous report (arXiv:1703.06739; to appear in Phys. Rev. Lett.), we have prese…

2018-02-16abs ↗pdf ↗

NEST optimizes deep learning training by placing devices efficiently across networks and memory.

problem Inefficient device placement in distributed deep learning leads to high communication and memory overhead.
method NEST uses network-, compute-, and memory-aware dynamic programming to optimize device placement.
result NEST achieves up to 2.43 times higher throughput and better memory efficiency.

We parallelize backpropagation for deep learning models, achieving significant speedups.

problem Sequential dependency in backpropagation limits scalability on parallel systems.
method Reformulated backpropagation as a scan operation, using Blelloch scan algorithm.
result Up to 2.75x speedup on overall training time and 108x on backward pass.

Predictability enables efficient parallelization of nonlinear models.

problem Understanding which nonlinear state space models can be efficiently parallelized.
method Established a relationship between system dynamics and optimization problem conditioning, quantified by the largest Lyapunov exponent.
result Predictable systems can be evaluated in O((logT)2)O((\log T)^2) time, improving over conventional sequential approaches.