Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1.6%3.1%4.7%6.3% · Sep 199719922001200920182026
48 results for Teacher Forcing

ITF improves DSR but inflates curvature, while marginal likelihood reduces it, affecting QoIs.

problem Curvature mismatch between teacher forcing and marginal likelihood in chaotic dynamical systems.
method Comparing objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of AL-RNNs.
result Curvature inflation by ITF and reduction by marginal likelihood affect dynamical quantities of interest.

Attention forcing improves sequence-to-sequence model training stability.

problem Training auto-regressive sequence-to-sequence models with attention mechanism is challenging.
method Attention forcing guides the model with generated output history and reference attention.
result Attention forcing trains models to recover from mistakes without requiring a schedule or classifier.

A new method for training generative models with sparse supervision.

problem Training deep generative models with sparse and varying supervision.
method Caffeinated Wake-Sleep (CWS) method, combining reweighted wake-sleep and teacher-forcing.
result The CWS method is robust to variable length supervision and performs well on various datasets.

Transformers solve parity problems efficiently with step-by-step reasoning.

problem Training transformers to solve complex, recursive problems like parity.
method Training a one-layer transformer to solve kk-parity, incorporating intermediate parities into the loss function, and using teacher forcing or augmented data.
result Transformers can learn parity in one gradient update with intermediate supervision or self-consistency checks.

Study finds exposure bias distortion is limited and not incremental in open-ended text generation.

problem Exposure bias in auto-regressive language models causing incremental distortion.
method Proposed metrics to quantify exposure bias impact, used ground-truth prefixes instead of model-generated prefixes.
result Exposure bias distortion is limited and not incremental during generation.

Enhances model compression with multi-teacher knowledge distillation.

problem Uncertainty evaluation and diverse teacher expertise in model deployment.
method Bayesian inference and teacher-informed prior with entropy-based weighting.
result Improved predictive accuracy and robust uncertainty quantification.

Aims to learn from multiple unpredictable teachers with minimal interaction.

problem Learning from multiple non-deterministic teachers with low interaction cost.
method Develops a framework and an active learning algorithm to estimate a distribution over policy space.
result Significantly reduces interaction with teachers without compromising performance.

Improved knowledge distillation using a teacher assistant to bridge the gap between student and teacher networks.

problem Large neural networks are hard to deploy on edge devices due to size constraints.
method Introduce multi-step knowledge distillation with an intermediate-sized teacher assistant.
result The proposed multi-step distillation method improves student network performance.

Student-teacher learning improves generalization with noisy inputs.

problem Transfer knowledge from clean inputs to noisy inputs.
method Analyzes student-teacher learning using deep linear networks and experiments with nonlinear networks.
result Three factors are vital for success: zero training loss, teacher knowledge, and feature decomposition.

Conditional T/S learning improves student model performance by selectively learning from teacher or ground truth.

problem Teacher's occasional wrong guidance leads to suboptimal student model performance.
method Proposes a conditional T/S learning scheme where the student selectively chooses between teacher and ground truth based on teacher correctness.
result The conditional learning achieves significant performance improvements over traditional T/S learning.

Paper tackles black-box machine teaching with cross-space models, proposing an active teacher model.

problem Teaching a learner with different feature representations and without full observation.
method Proposes an active teacher model that queries the learner to estimate its status and guide faster convergence.
result Active teacher model achieves faster convergence rate than traditional passive learning.

AC-Teach uses an ensemble of suboptimal teachers to improve exploration in RL.

problem Improving exploration efficiency in long-horizon tasks with sparse rewards.
method Bayesian Actor-Critic with an ensemble of suboptimal teachers.
result AC-Teach improves sample efficiency over baselines on various tasks.

A new teacher-class network method compresses DNNs by distributing knowledge to multiple student networks.

problem Overwhelming size of Deep Neural Networks (DNNs).
method Single teacher with multiple student networks, transferring knowledge to each student.
result The combined knowledge of the class of students achieves better performance and reduces parameters.

A new distillation method transfers channel information from teacher to student.

problem Transfer knowledge from teacher to student with fewer parameters and calculations.
method Channel Distillation (CD) and Guided Knowledge Distillation (GKD) with loss decay.
result Achieved 27.68% top-1 error on ImageNet with ResNet18, outperforming state-of-the-art methods.

Distilled models often fail to match teacher models, despite improving generalization.

problem The discrepancy between teacher and student predictive distributions remains large.
method Investigated the optimization difficulties and dataset details affecting student performance.
result Optimizing for matching the teacher does not always lead to better generalization.

Self distillation boosts CNN accuracy without increasing model size.

problem Improving CNN accuracy in resource-limited domains.
method Divide and compress knowledge within the network structure.
result Average accuracy improvement of 2.65% across various networks.

Random feature models can outperform a weak teacher with early stopping.

problem Generalization from a weak to a strong model in random feature networks.
method Random feature models, early stopping, proving weak-to-strong generalization.
result Random feature models can outperform a weak teacher with early stopping.

This paper investigates teacher hacking during language model distillation and proposes methods to mitigate it.

problem Teacher hacking during language model distillation, leading to suboptimal performance.
method A controlled experimental setup involving an oracle LM, teacher LM, and student LM, using fixed offline or online data generation techniques.
result Data diversity is the key factor in preventing teacher hacking during distillation.

This paper proposes a method to embed teacher knowledge into a student network without increasing parameters.

problem The need for portable neural networks on mobile devices with limited resources.
method Feature embedding approach to distill knowledge from a teacher network to a student network without introducing new parameters.
result The proposed method maintains the performance of the teacher network while significantly reducing computational and storage complexity.

Under-parameterized networks can either copy or average teacher weights, leading to universal optimal solutions.

problem Approximating a teacher network with an under-parameterized student network.
method Analyzing shallow neural networks with erf activation function and unitary teacher weights, proving copy-average configurations are critical points and finding the optimal solution.
result The optimal solution for under-parameterized networks has a universal structure, whether copying or averaging teacher neurons.

Study reveals how initial weights influence convergence in deep ReLU networks.

problem Understanding the dynamics and generalization of deep ReLU networks.
method Teacher-student setting, gradient analysis, and activation assumptions.
result Initial weights close to teacher nodes lead to faster convergence, and fan-out weights of other nodes converge to zero in over-parameterized cases.

Study shows how task similarity affects forgetting in teacher-student setup.

problem Catastrophic forgetting in continual learning.
method Extended teacher-student setup to multiple teachers, analyzing similarity between tasks.
result Task similarity, whether at readouts or features, influences forgetting and transfer.

KD can lead to student-teacher deviations that improve performance.

problem KD can lead to student-teacher deviations that may outperform the teacher.
method Characterized and explained the nature of student-teacher deviations through experiments and theory.
result KD can lead to improved generalization by exaggerating the implicit bias of gradient descent.

Role-wise data augmentation improves knowledge distillation effectiveness.

problem Existing knowledge distillation methods fail to utilize the full potential of teacher-student data interaction.
method Design and implement data augmentation agents with distinct roles for teacher and student.
result Specially tailored data points enhance the demonstration of teacher's knowledge to the student.

A new KD layer lets student models learn and apply teacher knowledge explicitly.

problem Implicit action of traditional KD on student's feature transform limits its use in intermediate layers.
method Proposes a learnable KD layer that explicitly embeds teacher's knowledge in feature transform.
result Improves KD with two abilities: leveraging teacher's knowledge and feeding forward knowledge deeper.

Improved learning to reweight using deep interactions between student and teacher models.

problem Limitation of existing learning to reweight methods in utilizing student model's internal states.
method Proposes an algorithm that uses the student model's internal states to the teacher model, which returns adaptive weights to enhance student model training.
result Significant improvement over previous methods in image classification and neural machine translation experiments.

The paper reveals three mechanisms for weak-to-strong generalization.

problem Understanding the mechanisms behind weak-to-strong generalization in imperfect labeling scenarios.
method Theoretical analysis of simple models including ridge regression and weighted ridge regression, and a nonlinear multi-index setting.
result A student model can compensate for a teacher's under-regularization and achieve lower test error.

Knowledge flow transfers knowledge from multiple teachers to a student net.

problem Choosing and initializing deep nets for new tasks is unclear and inefficient.
method Develops a method to move 'knowledge' from multiple deep nets (teachers) to a new net (student) without dependency on teachers.
result The student net outperforms fine-tuning and other methods on various tasks.

This work proposes a student-teacher network for predicting hospital admission locations.

problem Accurate prediction of hospital admission locations to optimize resource allocation.
method Reinforcement learning approach where a teacher network selects data batches for a student network.
result The approach outperforms state-of-the-art methods on tabular data and image recognition.

The paper analyzes knowledge distillation in wide neural networks, providing theoretical insights and practical implications.

problem Lack of theoretical understanding of knowledge distillation in wide neural networks.
method Theoretical analysis of knowledge distillation in a linearized model of a wide neural network, introducing a metric of task training difficulty.
result For a perfect teacher, a high ratio of teacher's soft labels can be beneficial. For imperfect teacher, hard labels can correct wrong predictions.

The study analyzes multi-class teacher-student perceptron performance and generalization errors.

problem Analyzing multi-class classification with the teacher-student perceptron.
method Deriving asymptotic expressions for Bayes-optimal and empirical risk minimization (ERM) generalization errors.
result Regularised cross-entropy minimization yields close-to-optimal accuracy for multi-class classification.

Improves learning efficiency for active sequential learners.

problem Optimizing training data for sequential learners who actively choose their queries.
method Formulated as a Markov decision process, addressing both teaching and learning from a teacher.
result Planning teaching and learner's model of the teacher improve learning outcomes.

A new training method improves MLIPs for faster, lighter simulations.

problem High computational and memory costs of complex MLIPs for large-scale MD simulations.
method Teacher-student training framework using latent atomic energy knowledge.
result Lightweight student MLIPs achieve faster MD speeds and comparable accuracy to teachers.

Two-layer ReLU networks outperform kernel methods in teacher-student settings.

problem Understanding the excess risk of two-layer ReLU neural networks in teacher-student models.
method Investigated a two-phase training process for a student network, comparing it to kernel methods.
result The student network reaches near-global optimality and outperforms kernel methods in minimax optimal rate.

Mean Teacher improves semi-supervised learning by averaging model weights, outperforming Temporal Ensembling.

problem Improving semi-supervised learning performance with limited labeled data.
method Averaging model weights instead of label predictions, penalizing inconsistency with an exponential moving average target.
result Mean Teacher achieves 4.35% error rate on SVHN with 250 labels, outperforming Temporal Ensembling with 1000 labels.

Paper analyzes how neural networks learn from a teacher in a specific setting.

problem Understanding how two-layer ReLU neural networks learn from a teacher in a regression model.
method Used gradient descent with specific regularization and over-parameterization, combined with measure representation and sparse estimation.
result Student network can identify teacher network parameters with high probability via gradient descent.

Interactive teaching speeds up IRL learning with adaptive demonstrations.

problem Tackles the challenge of accelerating IRL learning with teacher assistance.
method Interactive teaching framework where a teacher adapts demonstrations based on the learner's policy.
result Teaching algorithms converge in the omniscient setting, speeding up learning.

A new method transfers adversarial robustness from teacher to student using feature distillation.

problem Adversarial robustness transfer across different models and tasks.
method Guided Adversarial Contrastive Distillation (GACD) with contrastive learning and sample reweighted estimation.
result GACD effectively transfers adversarial robustness from teacher to student, achieving comparable or better results.