SNGM improves large-batch training accuracy.
problem Improving generalization in large-batch training.
method Stochastic Normalized Gradient Descent with Momentum.
result SNGM achieves better test accuracy than MSGD and other large-batch methods.
Large batch training improves deep learning performance without needing warmup.
problem Slow convergence at early epochs in large batch training.
method Proposes CLARS algorithm and analyzes convergence rate.
result Proposed algorithm outperforms gradual warmup and state-of-the-art large-batch optimizers.
Large batch sizes reduce gradient variance in DP-SGD, improving privacy.
problem Understanding why large batch sizes work in DP-SGD.
method Decomposed total gradient variance into subsampling and noise-induced variances, proving batch size independence in the limit.
result Large batch sizes reduce effective total gradient variance, improving privacy in DP-SGD.
Distributed implementations of mini-batch stochastic gradient descent (SGD) suffer from communication overheads, attributed to the high frequency of gradient updates inherent in small-batch training. Training with large batches can reduce these overheads; however, large batches can affect the convergence properties and…
DReg boosts large-batch SGD's generalization and convergence.
problem Large-batch SGD struggles with generalization in deep learning.
method DReg replicates a layer to encourage parameter diversity.
result DReg improves generalization and convergence with large-batch SGD.
LEGW improves large-batch training for both CNNs and RNNs.
problem Efficient large-batch training for deep neural networks.
method Linear-epoch gradual-warmup (LEGW) for large-batch training.
result Improved large-batch training for both CNNs and RNNs with Sqrt Scaling scheme.
AdaScale SGD adapts learning rates for large-batch training efficiently.
problem Adapting learning rates for large-batch training to balance speed-ups and model quality.
method Adaptive learning rate adaptation based on gradient variance.
result AdaScale achieves reliable speed-ups for a wide range of batch sizes without degrading model quality.
K-FAC doesn't improve large batch training efficiency.
problem Inefficiency of K-FAC in large batch size training.
method Empirical analysis of K-FAC and SGD on ResNet and AlexNet.
result K-FAC doesn't exhibit improved scalability to large batch sizes.
New framework optimizes deep learning training by deferring large batch sizes to late stages.
problem Optimizing batch size scheduling for deep learning training efficiency.
method Introduced the functional scaling law (FSL) framework to analyze and optimize batch size scheduling.
result Large batch sizes can be deferred to late training stages without sacrificing performance.
Standard optimizers perform as well as LARS and LAMB at large batch sizes.
problem Comparing optimizers for neural network training at large batch sizes.
method Used standard optimizers like Nesterov momentum and Adam to match or exceed LARS and LAMB results.
result Standard optimizers can match or exceed LARS and LAMB at large batch sizes.
SWAP uses large mini-batches to train DNNs faster with good generalization.
problem Training deep neural networks with small mini-batches is time-consuming.
method SWAP computes an approximate solution with large mini-batches and refines it by averaging weights of multiple parallel models.
result SWAP trains models as well as small-batch training but in significantly less time.
SP-NGD improves deep learning models' generalization with large mini-batch sizes.
problem Worse generalization performance with large mini-batch sizes in deep learning.
method SP-NGD, a natural gradient descent approach for large-scale deep learning.
result SP-NGD achieves similar generalization performance to first-order methods with accelerated convergence and negligible overhead.
This paper explores the limits of large batch sizes in deep learning.
problem Understanding the optimal batch size for deep learning models.
method Detailed numerical optimization and experimental analysis of large batch sizes.
result Improvement of 18% top-1 test accuracy with an optimized batch size recipe.
The study improves generalization in large-batch training by adding structured covariance noise to gradients.
problem Improving generalization in large-batch training while maintaining optimal convergence.
method Adding covariance noise to the gradients to improve generalization performance.
result The method improves generalization performance without degrading optimization performance and training duration.
Mini-batch EM algorithm speeds up convergence for large datasets.
problem Efficiently processing large datasets in latent variable models.
method Proposes mini-batch version of Stochastic Approximation EM algorithm for exponential models.
result Converges under classical conditions with mini-batch sampling.
Large batch sizes improve model performance but have limits; we found a simple predictor.
problem Understanding the limits of large batch sizes across different domains.
method Empirical model using gradient noise scale to predict optimal batch size.
result Gradient noise scale predicts the largest useful batch size across various domains.
New LAMB optimizer reduces BERT training time from 3 days to 76 minutes.
problem Training large deep neural networks on massive datasets is computationally challenging.
method Developed a new layerwise adaptive large batch optimization technique called LAMB.
result LAMB reduces BERT training time from 3 days to 76 minutes using very large batch sizes.
Large batch training with DP-SGD reduces model performance due to implicit bias.
problem Large batch training with DP-SGD reduces model performance.
method The study analyzes the phenomenon of implicit bias in Noisy-SGD (DP-SGD without clipping) and its theoretical solutions for linear models.
result The implicit bias in large batch training with DP-SGD is amplified by additional noise, similar to SGD.
Matching pursuit (MP) methods are a promising class of feature construction algorithms for value function approximation. Yet existing MP methods require creating a pool of potential features, mandating expert knowledge or enumeration of a large feature pool, both of which hinder scalability. This paper introduces batch…
DAL improves active learning for neural networks with large batch sizes.
problem Efficiently choosing examples to label for neural networks with large batch sizes.
method DAL treats active learning as a binary classification task to make labeled and unlabeled sets indistinguishable.
result DAL performs on par with state-of-the-art methods in medium and large query batch sizes.
Theoretical justification for large batch SGD converging to flatter minima.
problem Understanding the convergence of large batch SGD to flatter minima.
method Theoretical analysis of SGD in finite-time and asymptotic regimes.
result SGD tends to converge to flatter minima regardless of batch size, but with different rates.
ABS dynamically adjusts batch size based on policy stability, improving RL performance.
problem Diminishing returns with large batch sizes in RL due to non-stationary data.
method Adaptive Batch Scaling (ABS) with Behavioral Divergence metric.
result Larger batch sizes can improve RL performance, contrary to conventional wisdom.
Small-GAN speeds up GAN training by using coresets.
problem Slowness and high memory usage in GAN training with large batches.
method Draw a large batch of samples from the prior, compress using Coreset-selection, and use cached Inception activations for random projection.
result Significantly reduces training time and memory usage for modern GAN variants.
AdAdaGrad optimizes batch sizes for deep learning models, reducing the generalization gap.
problem The generalization gap between large-batch and small-batch training in deep learning.
method AdAdaGrad introduces adaptive batch size strategies derived from adaptive sampling methods.
result AdAdaGradNorm converges to a first-order stationary point with a rate of O(1/K) in K iterations.
Noise in SGD helps deep nets generalize better, even with smaller batch sizes.
problem The generalization benefit of using noise in SGD over large batch sizes.
method Carefully designed experiments and rigorous hyperparameter sweeps on various models.
result Small or moderately large batch sizes outperform very large batches on test sets.
Faster convergence and handling larger mini-batches for deep neural networks.
problem Generalization gap in large-scale distributed training of deep neural networks.
method Second-order optimization using Kronecker-factored approximate curvature.
result Achieved 75% Top-1 validation accuracy with mini-batch size of 131,072 in 978 iterations.
A new framework scales active search for large datasets.
problem Scaling active search for large, high-dimensional data sets.
method Hierarchical Batch Bandit Search (HBBS) framework.
result HBBS improves performance and scalability for batch search.
A new method improves active learning for large batch sizes.
problem Challenges in scaling Bayesian active learning to large batch sizes.
method Derives Partial Batch Label Sampling (ParBaLS) for EPIG algorithm.
result ParBaLS EPIG outperforms top-B selection and BatchBALD. Large batch size training of Neural Networks has been shown to incur accuracy loss when trained with the current methods. The exact underlying reasons for this are still not completely understood. Here, we study large batch size training through the lens of the Hessian operator and robust optimization. In particular, w…
Mini-batch stochastic gradient methods (SGD) are state of the art for distributed training of deep neural networks. Drastic increases in the mini-batch sizes have lead to key efficiency and scalability gains in recent years. However, progress faces a major roadblock, as models trained with large batches often do not ge…
Optimizes query routing to LLMs under cost and resource constraints.
problem Non-uniform or adversarial batching in per-query routing methods leads to cost inefficiency.
method Batch-level, resource-aware routing framework that jointly optimizes model assignment for each batch.
result Robust routing framework improves accuracy by 1-14% over non-robust methods.
Empirical study on SGD hyperparameters and adversarial robustness.
problem Effect of SGD hyperparameters on adversarial robustness and generalization.
method Empirical observation of learning rate, batch size, and momentum effects on adversarial robustness and generalization.
result Constant learning rate to batch size ratio leads to good generalization and almost constant adversarial robustness.
FRN layer eliminates batch dependence in deep learning, improving performance across various tasks.
problem Batch Normalization's dependency on mini-batch elements can degrade performance for small batches.
method Filter Response Normalization (FRN) operates independently on each activation channel of each batch element.
result FRN layer outperforms BN and other alternatives in various settings for all batch sizes.
Adaptive batch size schedules improve language model training efficiency and generalization.
problem Dilemma of choosing batch sizes in large-scale model training.
method General-purpose adaptive batch size schedules compatible with data and model parallelism.
result Adaptive batch size schedules outperform constant batch sizes and heuristic warmup schedules.
Accelerates BERT pretraining from 3 days to 54 minutes.
problem Long training time of BERT due to large mini-batch sizes.
method LANS method and learning rate scheduler for large mini-batch training.
result Achieved fastest BERT training time of 54 minutes.
Analysis of SGD+M convergence rates in high dimensions with batch size considerations.
problem Understanding convergence rates of SGD+M in high-dimensional settings.
method Analyzing the dynamics of SGD+M on least squares problems with large batch sizes and dimensions.
result Identifies the implicit conditioning ratio (ICR) that regulates SGD+M's acceleration and convergence rates.
GNC smooths loss function for large-batch SGD, improving generalization.
problem Extremely large-batch SGD leads to poor generalization and converges to sharp minima.
method Gradient noise convolution (GNC) smooths loss function by convolving gradient noise with the loss function.
result GNC achieves state-of-the-art generalization performance for large-scale deep neural networks.
Study shows how algorithmic choices affect optimal batch sizes in neural networks.
problem Understanding how batch size impacts neural network training efficiency.
method Experiments and analysis of a simple quadratic model to study algorithmic choices.
result Preconditioned optimizers like Adam and K-FAC allow larger batch sizes before diminishing returns.
Bayesian batch active learning approximates model parameters efficiently.
problem High label acquisition cost for large-scale supervised models.
method Sparse subset approximation using Bayesian active learning.
result Efficient active learning at scale with diverse batches.
Proposes qPO, a new acquisition strategy for batched Bayesian optimization that maximizes the probability of including the optimum.
problem Efficiently identifying top-performing compounds from a large chemical library.
method qPO (multipoint Probability of Optimality) acquisition strategy that maximizes the probability of including the true optimum.
result Empirical evidence shows that qPO is competitive with and complements other state-of-the-art methods in batched Bayesian optimization.
Momentum affects optimization differently at small vs large batch sizes near instability.
problem Understanding how momentum impacts optimization near the edge of stability.
method Demonstrated through batch-size dependent behavior of SGD with momentum.
result Momentum operates in two distinct regimes: amplifying stochastic fluctuations at small batch sizes and stabilizing at large batch sizes.
New method improves model accuracy in Byzantine-robust distributed learning by optimizing batch size.
problem Reduces model accuracy drop due to large variance of stochastic gradients in Byzantine-robust distributed learning.
method Proposes ByzSGDnm, a novel BRDL method that uses normalized momentum to mitigate accuracy drop in large batch sizes.
result The optimal batch size increases with the fraction of Byzantine workers, leading to better model accuracy under Byzantine attacks.
Seesaw optimizes training by balancing learning rate and batch size, accelerating model pretraining.
problem Optimizing training efficiency for large language models with adaptive optimizers.
method Develops a principled framework for batch-size scheduling, introducing Seesaw which multiplies learning rate by 1/√2 and doubles batch size.
result Empirically, Seesaw reduces wall-clock time by approximately 36% compared to cosine decay, matching theoretical limits.
Large batch sizes don't improve training time for most models.
problem The inefficiency of large batch sizes in stochastic gradient descent.
method Empirical analysis of network training across various architectures and domains.
result Increasing batch size beyond a certain point does not reduce training time for either train or test loss.
New attack recovers user-level information from large batch images.
problem Recovering private information from user-level gradients in distributed learning.
method Proposes a gradient inversion attack using a denoising diffusion model as a prior.
result Demonstrates recovery of realistic facial images and private attributes.
SogCLR uses small batch sizes for global contrastive learning, achieving similar performance to SimCLR.
problem Existing contrastive learning methods require large batch sizes or large feature dictionaries.
method SogCLR, a memory-efficient Stochastic Optimization algorithm for global contrastive learning.
result SogCLR with small batch sizes (e.g., 256) achieves similar performance to SimCLR with large batch sizes (e.g., 8192).
New analysis reveals batch size effects on stochastic conditional gradient methods.
problem Understanding the role of batch size in stochastic conditional gradient methods.
method Deriving a new analysis focusing on momentum-based stochastic conditional gradient algorithms (e.g., Scion).
result Increasing batch size initially improves optimization accuracy but can degrade performance beyond a critical threshold.
NLCG optimizes DNN training, especially with large mini-batches.
problem Improving convergence speed in large-scale DNN training.
method Stochastic Preconditioned Nonlinear Conjugate Gradient (SP-NLCG) algorithm.
result NLCG improves DNN training accuracy by over 10 percentage points at large mini-batch sizes.