Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

132264395527 · Jun 202019922001200920182026
48 results for Gradient Layers

A method predicts GNS of transformer layers using normalization layer norms.

problem Estimating gradient noise scale with minimal variance.
method Simultaneously compute per-example gradient norms and parameter gradients.
result Total GNS is predicted well by normalization layer GNS.

Gradient descent proves global convergence for 4-layer matrix factorization.

problem Global convergence of gradient descent on four-layer matrix factorization under random initialization.
method New techniques to show saddle-avoidance properties and extend eigenvalue theories.
result Polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization.

Gradient descent with logistic loss can make two-layer networks interpolate binary classification data.

problem Training two-layer networks for binary classification.
method Gradient descent with logistic loss applied to two-layer networks.
result Gradient descent can drive training loss to zero under certain conditions.

Enhances GAN training by adding a gradient layer to improve convergence.

problem Degeneration of convergence speed and limited representational power in GANs.
method Introduces a gradient layer to seek a descent direction in an infinite-dimensional space, bypassing local optima.
result Demonstrates faster convergence through numerical experiments.

Gradient descent achieves fast convergence for approximating functions with two-layer neural networks.

problem Approximating continuous functions with two-layer neural networks.
method Gradient descent combined with generic chaining technique from probability theory.
result Gradient descent yields an exponential convergence rate for two-layer neural networks without needing a large width relative to the number of data points.

Gradient descent balances layer magnitudes in deep neural networks without explicit regularization.

problem Balancing magnitudes across layers in deep neural networks.
method Gradient descent with infinitesimal step size enforces layer magnitude balance.
result Gradient descent automatically balances layer magnitudes without explicit regularization.

New Banach spaces for ReLU networks enable better function approximation and gradient dynamics analysis.

problem Function approximation and gradient dynamics in multi-layer ReLU networks.
method Developed Banach spaces for ReLU networks, defined new function representations, and analyzed gradient flow dynamics.
result Gradient flow dynamics of the new representation is the continuous analog of gradient descent for ReLU networks.

NovoGrad improves deep learning training with adaptive moments and layer-wise normalization.

problem Training deep neural networks efficiently and effectively.
method Layer-wise adaptive moments with gradient normalization and decoupled weight decay.
result NovoGrad outperforms well-tuned SGD with momentum and Adam/AdamW in various tasks.

Gradient descent learns two-layer networks well for classification problems.

problem Understanding the performance of gradient descent on classification problems.
method Refined convergence analysis of gradient descent for two-layer networks with smooth activations.
result Gradient descent can learn less over-parameterized networks for classification problems.

Proposes mGBDTs for learning hierarchical representations in gradient boosting decision trees.

problem Inability of gradient boosting decision trees to learn hierarchical representations.
method Introduces multi-layered GBDT forest (mGBDTs) with explicit emphasis on hierarchical learning.
result Jointly trained mGBDTs can learn hierarchical representations effectively without backpropagation.

SGD's training dynamics align with Hessian and gradient spectra in high-dimensional classification tasks.

problem Understanding the spectra of Hessian and gradient matrices in high-dimensional classification tasks.
method Rigorous analysis of SGD dynamics and spectra of Hessian and gradient matrices.
result SGD trajectory and emergent outlier eigenspaces align with a common low-dimensional subspace in multi-class high-dimensional mixtures and neural networks.

Study proposes memory-efficient backpropagation for linear layers in neural networks.

problem Significant memory usage in backpropagation through linear layers in neural networks.
method Randomized matrix multiplications to reduce memory usage with a moderate decrease in test accuracy.
result Demonstrated benefits of the proposed method on fine-tuning pre-trained models.

Proposes a method to balance tasks in multitask learning with a single gradient step update.

problem Balancing tasks in multitask learning to avoid imbalance.
method Gradient-based meta-learning to balance tasks at the gradient level, training shared and task-specific layers separately.
result Achieves state-of-the-art performance on various multitask computer vision problems.

LAGS-SGD optimizes deep learning training by sparsifying gradients layer-wise.

problem Reduces long training times in large deep neural networks with distributed S-SGD.
method Layer-wise adaptive gradient sparsification combined with S-SGD.
result LAGS-SGD achieves convergence guarantees and outperforms vanilla S-SGD.

Gradient descent learns ReLU networks with Gaussian inputs and noisy outputs.

problem Learning one-hidden-layer ReLU networks with Gaussian inputs and noisy outputs.
method Gradient descent with tensor initialization for empirical risk minimization.
result Gradient descent converges to ground-truth parameters at a linear rate up to statistical error.

Gradient descent converges to minimum Bayes risk for two-layer ReLU networks in mean field regime.

problem Training two-layer ReLU networks using gradient descent in the mean field regime.
method Describes a condition for convergence to minimum Bayes risk, extending previous results to ReLU-activated networks.
result The condition for convergence does not depend on initialization and concerns weak convergence of network realization.

Gradient descent proves global convergence for deep networks with a single wide layer.

problem Proving global convergence of gradient descent for deep ReLU networks.
method Simplified proof using a single wide layer, leveraging ReLU's Lipschitz property.
result Gradient descent converges globally for networks with a single wide layer.

Optimizes neural networks' last layer with closed-form solutions.

problem Optimizing neural networks' last layer with stochastic gradient descent.
method Adapting closed-form last layer optimization for stochastic gradient descent, alternating between backbone and last layer updates.
result The method converges to optimal solutions and outperforms standard SGD and Adam in regression tasks.

AdaLoss optimizes adaptive learning rates for efficient convergence in various models.

problem Efficiently optimizing adaptive learning rates for gradient descent methods.
method AdaLoss uses loss function information to dynamically adjust step sizes.
result AdaLoss achieves linear convergence in linear regression and robust global convergence in neural networks.

WarpGrad efficiently learns preconditioning matrices for gradient descent across task distributions.

problem Learning efficient update rules for rapid new task learning.
method Interleaves warp-layers between task-learner layers to meta-learn preconditioning matrices.
result WarpGrad scales to large meta-learning problems and improves across various learning settings.

Gradient descent struggles with high-dimensional data fitting.

problem Gradient descent struggles with high-dimensional data fitting.
method Gradient descent training of a two-layer neural network on empirical or population risk.
result Gradient descent training may not decrease population risk faster than t4/(d2)t^{-4/(d-2)} under mean field scaling.

This paper analyzes how training data can be leaked from gradients in neural networks and proposes a metric for measuring model security.

problem Training data leakage from gradients in neural networks for image classification.
method Formulated the problem as an optimisation problem for each layer, involving weights, gradients, and constraints from preceding layers.
result Attributed training data leakage to the architecture of the deep network and proposed a metric for measuring model security.

New algorithm trains deep neural networks with adaptive learning rates.

problem Inconsistent gradient magnitudes across layers in SGD.
method Back-matching propagation with approximations for layer-wise adaptive learning rates.
result Achieves favorable results over standard SGD in training deep neural networks.

Gradient descent optimizes neural networks and random features similarly, achieving zero loss fast.

problem Optimizing two-layer neural networks and random feature models under gradient descent.
method Comprehensive analysis of gradient descent dynamics, considering various network widths and data sizes.
result Gradient descent achieves zero training loss exponentially fast in the over-parametrized regime.

Automatically learns flexible symmetry constraints in neural networks using gradients.

problem Fixed hard constraints on neural network functions that cannot be adapted.
method Improves parameterisations of soft equivariance and optimizes marginal likelihood using differentiable Laplace approximations.
result Achieves equivalent or improved performance on image classification tasks compared to baselines with hard-coded symmetry.

A scalable framework for gradient boosting using TensorFlow.

problem Training gradient boosted trees efficiently on large datasets.
method Distributed training architecture, automatic loss differentiation, layer-by-layer boosting, multi-class handling, regularization.
result Faster prediction and smaller ensembles compared to traditional methods.

We establish the first benchmark for federated learning with differential privacy in ASR, achieving competitive performance.

problem Challenges in training large transformer models for ASR in federated learning with differential privacy.
method Per-layer clipping and layer-wise gradient normalization to mitigate gradient heterogeneity.
result FL with DP is viable in ASR with strong privacy guarantees, achieving competitive performance.

Transformers learn to perform logistic regression in-context.

problem Understanding how transformers learn to perform specific tasks in-context.
method Constructed multi-layer transformers that perform in-context logistic regression through normalized gradient descent.
result Transformers can be trained to perform in-context logistic regression effectively.

Gradient oversmoothing and expansion hinder deep GNN training, solved with normalization.

problem Gradient oversmoothing and expansion prevent deep GNN training.
method Proposed normalization method to constrain the Lipschitz bound of each layer.
result Residual GNNs with hundreds of layers can be efficiently trained with the proposed normalization.

Single wide layer followed by a pyramidal structure ensures global convergence in deep networks.

problem Ensuring global convergence in deep neural networks with limited width constraints.
method Proves that a single wide layer followed by a pyramidal structure guarantees global convergence for over-parameterized networks.
result Single wide layer of width NN suffices for global convergence in deep networks with constant-width remaining layers.

A new optimizer for deep learning improves accuracy and reduces training time.

problem Training deep neural networks for classification tasks.
method Hybrid Newton/Gradient Descent (NGD) method exploiting convexity of cross-entropy loss.
result Improves validation error and provides qualitative differences in hidden layer basis functions.

New metric shows how different regularization methods affect deep linear networks.

problem Understanding the training dynamics of deep linear networks.
method Introduced a new metric called layer imbalance to analyze training dynamics. Demonstrated behavior of different regularization methods and stochastic gradient descent.
result Different regularization methods behave similarly, leading to a flat minima.

Generalizes neural tangent kernel analysis for two-layer networks with noise and regularization.

problem Limitations of NTK analysis in deep learning practice.
method Generalized NTK analysis for two-layer neural networks with weight decay and gradient noise.
result Noisy gradient descent with weight decay exhibits 'kernel-like' behavior and converges linearly.

This paper extends stability and generalization analysis of GD for multi-layer NNs.

problem Understanding the generalization of multi-layer neural networks trained by GD.
method Comprehensive stability and generalization analysis of GD for multi-layer NNs, focusing on two-layer and three-layer networks.
result Derives excess risk rates of O(1/n)O(1/\sqrt{n}) for GD in two-layer and three-layer NNs under specific conditions.

The paper extends mean field results to three-layer neural networks using SGD.

problem Understanding the dynamics of training three-layer neural networks with SGD.
method Extending mean field results from two-layer networks to three-layer networks with two hidden layers, using non-linear partial differential equations.
result The distributions of weights in the two hidden layers are independent.

Layer normalization placement affects training stability and warm-up stage necessity.

problem Training instability and the necessity of a learning rate warm-up stage in Transformers.
method Theoretical analysis and mean field theory to prove gradient behavior at initialization.
result Removing the warm-up stage for Pre-LN Transformers can achieve comparable results with less time and tuning.

Stochastic gradient methods converge for training wide PINNs.

problem Convergence of stochastic gradient descent in training over-parameterized PINNs.
method Established linear convergence of stochastic gradient descent/flow in training over-parameterized two-layer PINNs.
result Linear convergence with high probability for general activation functions.

A simple 2-layer linear network outperforms neural networks in learning sparse targets.

problem Learning sparse targets from a sparse input with gradient descent.
method A 2-layer linear network with fully connected input layer and sparse targets.
result The 2-layer linear network achieves a lower expected square loss than neural networks.

Paper characterizes gradient descent dynamics for neural networks with finite width.

problem Characterize gradient descent dynamics for multi-layer neural networks.
method Non-asymptotic state evolution theory for finite-width networks.
result Gradient descent dynamics provide precise distributional characterization.