Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

222445667889 · Jun 202019922001200920182026
48 results for Layer Training

Stanza separates convolutional and fully connected layers for faster deep learning training.

problem Heavy data transfer between workers and servers in distributed deep learning.
method Layer separation: most nodes train convolutional layers, others train fully connected layers only.
result Significant acceleration of training time (1.34x--13.9x) over current systems.

Layer rotation predicts model generalization, improving test accuracy by up to 30%.

problem Predicting model generalization in deep networks.
method Monitoring the cosine distance between layer weights and their initial values during training.
result Training procedures that maximize layer rotation consistently lead to better generalization performance.

Pre-training is crucial for learning deep neural networks. Most of existing pre-training methods train simple models (e.g., restricted Boltzmann machines) and then stack them layer by layer to form the deep structure. This layer-wise pre-training has found strong theoretical foundation and broad empirical support. Howe…

2015-06-07abs ↗pdf ↗

This paper presents two unsupervised learning layers (UL layers) for label-free video analysis: one for fully connected layers, and the other for convolutional ones. The proposed UL layers can play two roles: they can be the cost function layer for providing global training signal; meanwhile they can be added to any re…

2017-05-24abs ↗pdf ↗

Two-layer CNNs can overfit well if initialized correctly.

problem Understanding the conditions for benign overfitting in over-parameterized CNNs.
method Extending analysis to fully trainable two-layer CNNs, examining initialization scaling effects.
result Initialization scaling of the output layer is crucial; large scales lead to fixed output behavior, small scales to complex interactions.

Traditionally, when generative models of data are developed via deep architectures, greedy layer-wise pre-training is employed. In a well-trained model, the lower layer of the architecture models the data distribution conditional upon the hidden variables, while the higher layers model the hidden distribution prior. Bu…

2014-05-06abs ↗pdf ↗

This paper analyzes how normalization layers improve neural network training.

problem Improving generalization performance and training speed of neural networks.
method Global convergence analysis of two-layer neural networks with ReLU activations and Weight Normalization.
result Introduction of normalization layers changes the optimization landscape, enabling faster convergence.

Paper trains multi-layer SNNs using NormAD for spatio-temporal error backpropagation.

problem Training multi-layer SNNs with non-linear integrate-and-fire dynamics.
method Formulates training as optimization, uses NormAD for iterative synaptic weight update.
result Validated on 2- and 3-layer SNNs solving spike-based XOR and generic problems.

The study analyzes a three-layer neural network's training dynamics using a functional-space mean-field theory.

problem Understanding the training dynamics of partially-trained three-layer neural networks.
method Generalized mean-field theory to functional spaces, proving convergence and feature learning.
result The training loss of the model decays to zero at a linear rate in the L2L_2 regression setting.

Improved recommendation systems using multi-layer embeddings reduce model size while maintaining accuracy.

problem Improving model accuracy in recommendation systems while minimizing model size.
method Introducing a multi-layer embedding training (MLET) architecture that trains embeddings via a sequence of linear layers.
result Substantial advantages in model accuracy and memory footprint are achieved with reduced embedding dimensions.

Layer-wise training for deep linear networks achieves faster convergence with optimal learning rate.

problem Training deep neural networks is challenging; layer-wise training is proposed as an alternative.
method Layer-wise training using block coordinate gradient descent (BCGD) with orthogonal-like initialization.
result The optimal learning rate guarantees the fastest decrease in loss and is applicable without prior knowledge.

Designs a neural network to reduce training cost by mapping to higher dimensions.

problem High training cost in neural networks.
method Maps feature vectors to higher dimensional space, designs weight matrices to reduce cost, uses convex constraints.
result Reduces training cost as the number of layers increases, without cross-validation.

Study on rich regime training in deep learning, finding active parameters in bottom layers.

problem Understanding the practical success of deep learning models.
method Empirical study on rich regime training with benchmark datasets, re-initialization analysis, and probabilistic Layer-Wise Sparse SGD.
result Probabilistic Layer-Wise Sparse SGD matches vanilla SGD's generalization performance with improved efficiency.

New method makes neural networks more secure by protecting latent layers from adversarial attacks.

problem Vulnerability of latent layers in adversarially trained models to small perturbations.
method Latent Adversarial Training (LAT) and Latent Attack (LA) algorithms.
result Improves adversarial accuracy by 1-2% on MNIST, CIFAR-10, CIFAR-100 datasets.

This study analyzes how one-layer transformers learn regular language recognition tasks.

problem Understanding how one-layer transformers solve regular language recognition tasks like even pairs and parity check.
method Theoretical analysis of training dynamics and gradient descent for a one-layer transformer.
result A one-layer transformer can solve even pairs directly but needs CoT for parity check. Training phases show rapid growth in attention layer followed by logarithmic growth in linear layer.

Proposes a new architecture to speed up large-scale CNN training.

problem Time-consuming training process of large-scale CNNs.
method Bi-layered Parallel Training (BPT-CNN) architecture with outer-layer and inner-layer parallelism.
result Improves training performance of CNNs while maintaining accuracy.

Kernelized neural networks learn layers without backpropagation.

problem Learning deep feedforward architectures without backpropagation.
method Kernelize neural networks, present a greedy layer-wise training scheme.
result Kernelized networks trained layer-wise perform favorably compared to classical methods.

Modified RV-coefficient reveals how training affects neural network representations.

problem Understanding how training affects intermediate representations in convolutional neural networks.
method Experimented with modified RV-coefficient (RV2) to compare activation patterns in deep networks trained on varying amounts of data and layers.
result RV2 successfully recovered expected similarity patterns and provided interpretable similarity matrices.

Optimizes CNN architectures by analyzing receptive fields without training.

problem CNNs often have too many layers, making them resource-intensive and less effective.
method Layer-wise analysis of receptive fields to identify unproductive layers.
result Identifies and removes unproductive layers, improving CNN performance and efficiency.

Normalization layers improve the accuracy of Differentially Private training of deep neural networks.

problem Reduced accuracy in deep neural networks with Differentially Private training.
method Proposed a novel method for integrating batch normalization with Differentially Private Stochastic Gradient Descent (DPSGD) without additional privacy loss.
result Training deeper networks with better utility-privacy trade-off is possible.

Enhances GAN training by adding a gradient layer to improve convergence.

problem Degeneration of convergence speed and limited representational power in GANs.
method Introduces a gradient layer to seek a descent direction in an infinite-dimensional space, bypassing local optima.
result Demonstrates faster convergence through numerical experiments.

New algorithm trains deep neural networks with adaptive learning rates.

problem Inconsistent gradient magnitudes across layers in SGD.
method Back-matching propagation with approximations for layer-wise adaptive learning rates.
result Achieves favorable results over standard SGD in training deep neural networks.

Training state-of-the-art, deep neural networks is computationally expensive. One way to reduce the training time is to normalize the activities of the neurons. A recently introduced technique called batch normalization uses the distribution of the summed input to a neuron over a mini-batch of training cases to compute…

2016-07-21abs ↗pdf ↗

Backpropagation-free RL method trains layers using local signals.

problem Vanishing or exploding gradients in backpropagation-based RL.
method Local pairwise distance matching for layer-wise training without backpropagation.
result Backpropagation-free method achieves competitive performance and stability.

A new neural network initialization method is proposed for faster and more accurate training.

problem Efficient initialization for training multi-layer feedforward neural networks.
method Initialization based on Stein's identity, using eigenvectors of cross-moment matrix.
result The SteinGLM method is faster and more accurate than other initialization methods.

The paper introduces a multilevel initialization method for deep neural networks.

problem Training very deep neural networks with layer-parallel methods.
method Continuous interpretation of training as optimal control, using time-dependent ODEs for neural network discretization, and a refinement strategy across the time domain.
result The method creates deep networks with good initializations from coarser networks, reducing training time and providing regularization.

This paper shows how to train only the implicit layer of overparameterized implicit neural networks.

problem Understanding how the implicit layer contributes to the training of overparameterized implicit neural networks.
method Restricting training to only the implicit layer and analyzing the generalization error for ReLU-activated networks.
result Global convergence is guaranteed even if only the implicit layer is trained, and gradient flow with proper random initialization can achieve small generalization errors.

A new optimizer for deep learning improves accuracy and reduces training time.

problem Training deep neural networks for classification tasks.
method Hybrid Newton/Gradient Descent (NGD) method exploiting convexity of cross-entropy loss.
result Improves validation error and provides qualitative differences in hidden layer basis functions.

Proposes BN layers for neural networks on complex domains, improving training stability and accuracy.

problem Training stability and accuracy issues in neural networks on complex domains.
method Developed Riemannian batch normalization (BN) layers with connections to existing layers.
result Demonstrated improved performance on radar clutter classification, node classification, and action recognition.

When using deep, multi-layered architectures to build generative models of data, it is difficult to train all layers at once. We propose a layer-wise training procedure admitting a performance guarantee compared to the global optimum. It is based on an optimistic proxy of future performance, the best latent marginal. W…

2012-12-07abs ↗pdf ↗

Layered neural networks have greatly improved the performance of various applications including image processing, speech recognition, natural language processing, and bioinformatics. However, it is still difficult to discover or interpret knowledge from the inference provided by a layered neural network, since its inte…

2017-03-01abs ↗pdf ↗

Proposes QEP to mitigate quantization error propagation in layer-wise post-training quantization.

problem Growth of quantization errors across layers degrades performance, especially in low-bit regimes.
method Quantization Error Propagation (QEP) framework that explicitly propagates and compensates for quantization errors.
result QEP-enhanced layer-wise PTQ achieves substantially higher accuracy, especially in low-bit regimes.