Aims to learn from multiple unpredictable teachers with minimal interaction.
problem Learning from multiple non-deterministic teachers with low interaction cost.
method Develops a framework and an active learning algorithm to estimate a distribution over policy space.
result Significantly reduces interaction with teachers without compromising performance.
A new teacher-class network method compresses DNNs by distributing knowledge to multiple student networks.
problem Overwhelming size of Deep Neural Networks (DNNs).
method Single teacher with multiple student networks, transferring knowledge to each student.
result The combined knowledge of the class of students achieves better performance and reduces parameters.
Enhances model compression with multi-teacher knowledge distillation.
problem Uncertainty evaluation and diverse teacher expertise in model deployment.
method Bayesian inference and teacher-informed prior with entropy-based weighting.
result Improved predictive accuracy and robust uncertainty quantification.
Knowledge flow transfers knowledge from multiple teachers to a student net.
problem Choosing and initializing deep nets for new tasks is unclear and inefficient.
method Develops a method to move 'knowledge' from multiple deep nets (teachers) to a new net (student) without dependency on teachers.
result The student net outperforms fine-tuning and other methods on various tasks.
Study reveals how initial weights influence convergence in deep ReLU networks.
problem Understanding the dynamics and generalization of deep ReLU networks.
method Teacher-student setting, gradient analysis, and activation assumptions.
result Initial weights close to teacher nodes lead to faster convergence, and fan-out weights of other nodes converge to zero in over-parameterized cases.
ReOPD uses pre-collected teacher trajectories to distill knowledge from multi-turn interactions.
problem The cost of fully online on-policy distillation for multi-turn interactions.
method ReOPD, an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes, addressing the prefix trap and distribution shift.
result ReOPD preserves or improves OPD-level accuracy, uses zero tool calls, and is at least 4imes faster per training step. Robot learns from multiple teachers to efficiently achieve various motor skill outcomes.
problem Efficiently learning motor skills from multiple teachers and strategies.
method Hierarchical active decisions based on empirical evaluation of learning progress.
result Significantly more efficient learning and coherent strategy selection.
Study shows how task similarity affects forgetting in teacher-student setup.
problem Catastrophic forgetting in continual learning.
method Extended teacher-student setup to multiple teachers, analyzing similarity between tasks.
result Task similarity, whether at readouts or features, influences forgetting and transfer.
Self-distillation improves model performance by increasing teacher diversity and smoothing predictions.
problem Improving model generalization and performance through self-distillation.
method Interpreting self-distillation as MAP estimation and proposing instance-specific label smoothing.
result Self-distillation enhances model performance by increasing teacher diversity and smoothing predictions.
Framework transfers limited steering angle data across multiple weather conditions.
problem Limited labeled data for diverse weather conditions in sensorimotor control.
method Teacher-student learning paradigm with image-to-image translation network.
result Framework generalizes well across multiple weather conditions using limited labels.
Fastest video anomaly detection via teacher-student distillation.
problem Anomaly detection in video at high speed.
method Adversarial knowledge distillation from object-level teacher models.
result 7-62 times faster than state-of-the-art methods.
Paper distills word embeddings to reduce dimensionality without sacrificing accuracy.
problem Reducing neural model size for practical deployment.
method Teacher-student model-based embedding distillation with ensemble learning.
result Significant reduction in model size (80x faster and lighter) with minimal accuracy loss.
Student network specializes teacher nodes in deep ReLU networks.
problem Training deep ReLU networks with finite width and input dimension.
method Stochastic Gradient Descent (SGD) on over-realized student network trained from teacher network output.
result Each teacher node is specialized by at least one student node at the lowest layer under mild conditions.
A collaborative machine teaching method that improves learner performance with privacy and efficiency.
problem Improving learner performance with distributed teachers while maintaining privacy and scalability.
method Formulates collaborative teaching as a consensus and privacy-preserving optimization process to minimize teaching risk.
result The proposed method delivers significantly more accurate teaching results with high speed compared to non-collaborative MINLP-based super teaching.
Some machine learning applications involve training data that is sensitive, such as the medical histories of patients in a clinical trial. A model may inadvertently and implicitly store some of its training data; careful analysis of the model may therefore reveal sensitive information. To address this problem, we demon…
A new RL framework helps align RL solutions with existing practices.
problem Discrepancies between RL solutions and existing procedures.
method Student-teacher RL mechanism with a constraint on policy difference.
result Proves asymptotic optimality and effectiveness in multiple scenarios.
A new learning framework uses a teacher to guide learners through rounds of exposure to data.
problem Computational intractability of supervised learning.
method Curricular phases with a teacher guiding exposure to data.
result Simple learners can acquire complex concepts efficiently from examples.
ITF improves DSR but inflates curvature, while marginal likelihood reduces it, affecting QoIs.
problem Curvature mismatch between teacher forcing and marginal likelihood in chaotic dynamical systems.
method Comparing objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of AL-RNNs.
result Curvature inflation by ITF and reduction by marginal likelihood affect dynamical quantities of interest.
KD-Net transfers knowledge from multi-modal to mono-modal segmentation networks.
problem Limited acquisition of multiple imaging modalities in clinical settings.
method Generalized distillation framework adapted for mono-modal networks.
result The student network outperforms baseline mono-modal networks in brain tumor segmentation.
Yes, they do. This paper provides the first empirical demonstration that deep convolutional models really need to be both deep and convolutional, even when trained with methods such as distillation that allow small or shallow models of high accuracy to be trained. Although previous research showed that shallow feed-for…
Paper introduces a simulator-free approach to reinforcement learning policy distillation.
problem Learning multiplicity of cases corresponding to a given action in reinforcement learning.
method Generative adversarial approach to find multiple exemplars for each output class.
result Improves over state-of-the-art on data-free learning of student networks.
A new method trains a smaller model from a larger one without needing the actual training data.
problem Training a smaller model from a larger one without access to the training data.
method Synthesizes data impressions from the Teacher model to train the Student model.
result Zero-Shot Knowledge Distillation achieves competitive generalization performance.
NoNN compresses deep networks into distributed IoT modules with minimal communication.
problem Memory and communication constraints in IoT devices for deep learning inference.
method NoNN compresses a large pretrained network into disjoint, highly-compressed student modules, optimizing for memory and communication.
result NoNN achieves higher accuracy than baselines and similar to the teacher model with minimal communication.
Proposes RaT to mitigate bias in student-teacher estimation.
problem Systematic bias in teacher's predictions propagates to student model.
method Uses teacher to estimate residuals in student's predictions.
result RaT method reduces teacher bias effect and achieves optimal rate.
Paper introduces efficient uncertainty estimation in LLMs without multiple forward passes.
problem Accurate uncertainty quantification in LLMs remains challenging.
method Evidential Knowledge Distillation to create compact student models.
result Efficient uncertainty estimation achieved with single forward pass.
Unified scoring model improves efficiency and performance across multiple tasks.
problem Efficient and resource-efficient automated scoring for diverse tasks.
method Knowledge-distilled multi-task Mixture-of-Experts (MoE) approach.
result Comparable performance to task-specific models with significantly less storage and training resources.
Develops a generic approach for stable model distillation.
problem Stability of student models in model distillation when data variability is present.
method Generic approach based on central limit theorem for average loss, multiple testing framework.
result Demonstrates selection of a corpus size for consistent student models across different pseudo samples.
OKDDip uses diverse peers to improve online knowledge distillation.
problem Early saturation in group-based distillation.
method Two-level distillation with multiple auxiliary peers and a group leader, using attention-based aggregation weights.
result OKDDip consistently gives better performance than state-of-the-art approaches.
Improved knowledge distillation using a teacher assistant to bridge the gap between student and teacher networks.
problem Large neural networks are hard to deploy on edge devices due to size constraints.
method Introduce multi-step knowledge distillation with an intermediate-sized teacher assistant.
result The proposed multi-step distillation method improves student network performance.
Student-teacher learning improves generalization with noisy inputs.
problem Transfer knowledge from clean inputs to noisy inputs.
method Analyzes student-teacher learning using deep linear networks and experiments with nonlinear networks.
result Three factors are vital for success: zero training loss, teacher knowledge, and feature decomposition.
Conditional T/S learning improves student model performance by selectively learning from teacher or ground truth.
problem Teacher's occasional wrong guidance leads to suboptimal student model performance.
method Proposes a conditional T/S learning scheme where the student selectively chooses between teacher and ground truth based on teacher correctness.
result The conditional learning achieves significant performance improvements over traditional T/S learning.
Paper tackles black-box machine teaching with cross-space models, proposing an active teacher model.
problem Teaching a learner with different feature representations and without full observation.
method Proposes an active teacher model that queries the learner to estimate its status and guide faster convergence.
result Active teacher model achieves faster convergence rate than traditional passive learning.
AC-Teach uses an ensemble of suboptimal teachers to improve exploration in RL.
problem Improving exploration efficiency in long-horizon tasks with sparse rewards.
method Bayesian Actor-Critic with an ensemble of suboptimal teachers.
result AC-Teach improves sample efficiency over baselines on various tasks.
Subclass distillation improves small models by matching teacher's subclass probabilities.
problem Improving small models trained on limited data.
method Train a small model to match probabilities of subclasses invented by a large teacher model.
result Better small models trained on limited data.
A new distillation method transfers channel information from teacher to student.
problem Transfer knowledge from teacher to student with fewer parameters and calculations.
method Channel Distillation (CD) and Guided Knowledge Distillation (GKD) with loss decay.
result Achieved 27.68% top-1 error on ImageNet with ResNet18, outperforming state-of-the-art methods.
Teaches manifolds from teacher's structured data.
problem Learning a manifold from teacher's structured data.
method Extends existing approaches to learning from randomly sampled data points, considering structured data provided by a teacher.
result Demonstrations can significantly reduce data points needed for manifold learning.
Over-parametrization speeds up learning a single neuron model.
problem Understanding why over-parametrization accelerates learning in neural networks.
method Studied a simple model of a single teacher neuron with quadratic activation, showing how over-parametrization can lead to faster convergence.
result Over-parametrization helps gradient descent enter the neighborhood of a global optimal solution faster.
Distilled models often fail to match teacher models, despite improving generalization.
problem The discrepancy between teacher and student predictive distributions remains large.
method Investigated the optimization difficulties and dataset details affecting student performance.
result Optimizing for matching the teacher does not always lead to better generalization.
A new method transfers knowledge without data, matching teacher's predictions closely.
problem Lack of access to training data for knowledge transfer.
method Adversarial training to match teacher's predictions without data.
result Zero-shot student performs well on CIFAR10, improving state-of-the-art.
Random feature models can outperform a weak teacher with early stopping.
problem Generalization from a weak to a strong model in random feature networks.
method Random feature models, early stopping, proving weak-to-strong generalization.
result Random feature models can outperform a weak teacher with early stopping.
This paper investigates teacher hacking during language model distillation and proposes methods to mitigate it.
problem Teacher hacking during language model distillation, leading to suboptimal performance.
method A controlled experimental setup involving an oracle LM, teacher LM, and student LM, using fixed offline or online data generation techniques.
result Data diversity is the key factor in preventing teacher hacking during distillation.
This paper proposes a method to embed teacher knowledge into a student network without increasing parameters.
problem The need for portable neural networks on mobile devices with limited resources.
method Feature embedding approach to distill knowledge from a teacher network to a student network without introducing new parameters.
result The proposed method maintains the performance of the teacher network while significantly reducing computational and storage complexity.
Under-parameterized networks can either copy or average teacher weights, leading to universal optimal solutions.
problem Approximating a teacher network with an under-parameterized student network.
method Analyzing shallow neural networks with erf activation function and unitary teacher weights, proving copy-average configurations are critical points and finding the optimal solution.
result The optimal solution for under-parameterized networks has a universal structure, whether copying or averaging teacher neurons.
Paper proposes a method to combine multiple facial analysis models for better performance.
problem Facial analysis models from different sources have low transferability.
method Two-step process: 1) Auto-encoder for common embedding, 2) Distillation for lightweight model.
result Lightweight model outperforms state-of-the-art on 15 facial analysis tasks.
Estimates model performance from compute budget for distillation.
problem Risk mitigation in large-scale distillation.
method Distillation scaling law based on compute budget allocation.
result Maximizes student performance with compute-optimal allocation.
We generalise the problem of inverse reinforcement learning to multiple tasks, from multiple demonstrations. Each one may represent one expert trying to solve a different task, or as different experts trying to solve the same task. Our main contribution is to formalise the problem as statistical preference elicitation,…
BANs outperform teachers in computer vision and language modeling.
problem Improving model performance while reducing model size.
method Train students identically to their teachers using KD.
result BANs achieve state-of-the-art performance on CIFAR-10 and CIFAR-100 datasets.
KD can lead to student-teacher deviations that improve performance.
problem KD can lead to student-teacher deviations that may outperform the teacher.
method Characterized and explained the nature of student-teacher deviations through experiments and theory.
result KD can lead to improved generalization by exaggerating the implicit bias of gradient descent.