This paper improves student networks to be more robust against perturbations.
problem Improving robustness of lightweight student networks.
method The paper introduces a method to make student networks more confident and robust by leveraging teacher networks.
result The proposed method enhances the robustness of student networks without sacrificing accuracy.
Enhances student diversity in collaborative learning.
problem Student homogenization in large groups.
method Random routing, diverse feature sets, and random subgroup imitation.
result Significantly outperforms state-of-the-art approaches.
Developed predictive models for improving programming course performance.
problem Improving student performance in programming courses.
method Used M5P Decision Tree and Linear Regression Classifier on structured questionnaire data.
result Variable-based LRC model produced the best model with least evaluation metrics.
GritNet predicts student performance using deep learning.
problem Predicting student performance in online coursework.
method GritNet is a deep learning algorithm based on bidirectional long short term memory (BLSTM).
result GritNet outperforms logistic regression and improves predictions in the early stages of a course.
Predicts student performance in interactive online question pools using GNNs.
problem Predicting student performance in interactive online question pools with evolving knowledge.
method Proposes R^2GCN, a GNN model for heterogeneous networks to predict student performance.
result Achieves higher accuracy in student performance prediction than traditional methods.
Student-teacher learning improves generalization with noisy inputs.
problem Transfer knowledge from clean inputs to noisy inputs.
method Analyzes student-teacher learning using deep linear networks and experiments with nonlinear networks.
result Three factors are vital for success: zero training loss, teacher knowledge, and feature decomposition.
The Open University studies student online behavior in virtual learning environments.
problem Improving retention rates in online modules.
method GUHA and Markov chain-based analysis of student activity.
result Both methods are valid for modeling student activities.
Study examines how different assessment formats affect student learning in a data communications course.
problem Understanding how various assessment formats impact student learning outcomes.
method Comparing student learning outcomes across multiple assessment formats in a core data communications course at George Mason University.
result Collective assessment formats enhance student knowledge demonstration.
Improved ImageNet classification with semi-supervised learning.
problem Image classification with limited labeled data.
method Noisy Student Training: semi-supervised learning with noisy student models.
result 88.4% top-1 accuracy on ImageNet, 2.0% better than state-of-the-art.
A new teacher-class network method compresses DNNs by distributing knowledge to multiple student networks.
problem Overwhelming size of Deep Neural Networks (DNNs).
method Single teacher with multiple student networks, transferring knowledge to each student.
result The combined knowledge of the class of students achieves better performance and reduces parameters.
Predicts student outcomes in real-time using domain adaptation.
problem Real-time student performance prediction in online courses.
method GritNet architecture with unsupervised domain adaptation.
result GritNet enhances real-time predictions, especially in early weeks.
Paper presents deep learning and ML for automated student performance estimation.
problem Evaluation of students' performance during the pandemic.
method In-depth analysis of deep learning and machine learning approaches.
result Better performance across different prediction tasks with fully data-driven approach.
Improved learning to reweight using deep interactions between student and teacher models.
problem Limitation of existing learning to reweight methods in utilizing student model's internal states.
method Proposes an algorithm that uses the student model's internal states to the teacher model, which returns adaptive weights to enhance student model training.
result Significant improvement over previous methods in image classification and neural machine translation experiments.
Study analyzes online student behavior patterns using log data.
problem Understanding and optimizing student learning in online educational systems.
method Non-negative matrix factorization techniques for soft clustering.
result Behavioral changes of individual students and the system over time.
Personalized education at scale improves student outcomes.
problem Traditional personalized education is expensive and inequitable.
method Adapting educational presentations using RL, semi-supervised learning, NLP, and CV.
result Personalized education at scale can improve student outcomes.
Modeling student behaviors and multiple predictions for early intervention.
problem Predicting student outcomes and interactions among multiple tasks.
method Proposes a variant of LSTM and soft-attention mechanism for heterogeneous behaviors, and co-attention mechanism for task interactions.
result Demonstrated effectiveness in predicting student outcomes and interactions.
Estimates model performance from compute budget for distillation.
problem Risk mitigation in large-scale distillation.
method Distillation scaling law based on compute budget allocation.
result Maximizes student performance with compute-optimal allocation.
Method teaches students without teachers, estimating true labels from crowdsourcing.
problem Teaching without access to true labels.
method Apply crowdsourcing techniques to estimate true labels and student models for iterative teaching.
result Teaching performance is particularly effective for low-level students.
A new KD layer lets student models learn and apply teacher knowledge explicitly.
problem Implicit action of traditional KD on student's feature transform limits its use in intermediate layers.
method Proposes a learnable KD layer that explicitly embeds teacher's knowledge in feature transform.
result Improves KD with two abilities: leveraging teacher's knowledge and feeding forward knowledge deeper.
Teaches AI models to learn effectively through dynamic loss functions.
problem Optimizing machine learning models' performance through dynamic loss functions.
method Develops a method for a teacher model to dynamically output loss functions for a student model, enabling gradient-based optimization.
result Significantly improves the performance of various student models in real-world tasks.
A new method for student-initiated action advice using novelty detection.
problem Exploration and sample inefficiency in RL, especially with teacher absence.
method Random Network Distillation (RND) to measure advice novelty, updates only for advised states.
result Significant performance improvement over state-of-the-art methods, especially in challenging scenarios.
Each year, roughly 30% of first-year students at US baccalaureate institutions do not return for their second year and over $9 billion is spent educating these students. Yet, little quantitative research has analyzed the causes and possible remedies for student attrition. Here, we describe initial efforts to model stud…
Conditional T/S learning improves student model performance by selectively learning from teacher or ground truth.
problem Teacher's occasional wrong guidance leads to suboptimal student model performance.
method Proposes a conditional T/S learning scheme where the student selectively chooses between teacher and ground truth based on teacher correctness.
result The conditional learning achieves significant performance improvements over traditional T/S learning.
Paper tackles zero-shot code feedback using rubric sampling with deep learning.
problem Lack of historical data for supervised learning in introductory programming assignments.
method Human-in-the-loop rubric sampling with deep learning inference.
result Autonomous feedback for first students is more accurate than data-hungry algorithms and approaches human level fidelity.
Two-layer ReLU networks outperform kernel methods in teacher-student settings.
problem Understanding the excess risk of two-layer ReLU neural networks in teacher-student models.
method Investigated a two-phase training process for a student network, comparing it to kernel methods.
result The student network reaches near-global optimality and outperforms kernel methods in minimax optimal rate.
The paper develops a student performance prediction model using ensemble methods.
problem Improving student performance prediction in high school courses.
method Developed a multilabel classification model using SVM, RF, KNN, and MLP. Improved performance with LP transformation and partitioning schemes.
result The model achieved better performance than binary relevance and classifier chains.
Proposes KTAN for better training of student networks with both intermediate representations and probability distributions.
problem Reduces large computation and storage cost of deep networks by transferring generalization ability.
method Holistically considers intermediate representations and probability distributions; uses a Teacher-to-Student layer and adversarial learning.
result Significantly improves performance of student networks on image classification and object detection tasks.
A new method transfers knowledge without data, matching teacher's predictions closely.
problem Lack of access to training data for knowledge transfer.
method Adversarial training to match teacher's predictions without data.
result Zero-shot student performs well on CIFAR10, improving state-of-the-art.
SiamJEPA uses Siamese student encoders to improve JEPA-based representation learning.
problem Improving self-supervised representation learning in JEPA models.
method Proposes SiamJEPA with masked Siamese student encoders and EMA teacher network.
result Siamese student encoders improve representation separability and learning speed.
This work proposes a student-teacher network for predicting hospital admission locations.
problem Accurate prediction of hospital admission locations to optimize resource allocation.
method Reinforcement learning approach where a teacher network selects data batches for a student network.
result The approach outperforms state-of-the-art methods on tabular data and image recognition.
Paper generalizes teacher-student model for realistic data.
problem Capturing learning curves for realistic datasets.
method Introduces a Gaussian covariate generalization of the teacher-student model.
result Generalized model captures learning curves for various realistic data sets.
New method provides fine-grained feedback on interactive student programs.
problem Time-consuming manual grading of interactive student programs.
method Meta-exploration approach using reinforcement learning.
result 94.3% accuracy in providing fine-grained feedback.
A new training method improves MLIPs for faster, lighter simulations.
problem High computational and memory costs of complex MLIPs for large-scale MD simulations.
method Teacher-student training framework using latent atomic energy knowledge.
result Lightweight student MLIPs achieve faster MD speeds and comparable accuracy to teachers.
Dual Student separates the teacher from the student in SSL, improving performance.
problem Performance bottleneck caused by coupled teacher in consistency-based SSL methods.
method Introduces Dual Student, replacing the teacher with another student and defining a stabilization constraint.
result Significant improvement in classification performance on SSL benchmarks.
Optimal learning paths designed for E-learning systems using reinforcement learning.
problem Designing optimal learning paths for E-learning systems.
method Developed a hierarchical skill model and a proficiency level model, applied reinforcement learning to find the optimal learning strategy.
result Demonstrated the effectiveness of the proposed framework via numerical experiments.
Modeling student course choices using latent variables.
problem Understanding student enrollment patterns in large universities.
method Probabilistic approach based on multilabel classification and mixture models.
result Demonstrated the model's ability to infer student interests guiding enrollment decisions.
Study shows a positive effect of mindset interventions on student performance, with significant heterogeneity.
problem Assessing the effect of student mindset interventions and identifying factors that moderate this effect.
method Used machine learning techniques to address treatment overlap, balance, and imputation of conditional average treatment effects.
result The mindset intervention has a positive average effect of 0.26, with heterogeneity ranging from 0.1 to 0.4, moderated by school-level factors.
Proposes grade-aware course recommendation methods to improve student GPA.
problem Helping students select courses that lead to timely graduation and good grades.
method Two approaches: ranking courses by expected GPA impact and combining grade predictions with course recommendations.
result Grade-aware methods recommend courses leading to better student performance.
A new system recommends specific knowledge concepts in MOOCs based on student interests.
problem MOOCs recommend courses but ignore specific knowledge concepts interests.
method End-to-end graph neural network (ACKRec) that combines content and context information.
result ACKRec effectively recommends knowledge concepts to MOOC students.
Study shows how task similarity affects forgetting in teacher-student setup.
problem Catastrophic forgetting in continual learning.
method Extended teacher-student setup to multiple teachers, analyzing similarity between tasks.
result Task similarity, whether at readouts or features, influences forgetting and transfer.
A student classifier learns from feature visualizations of a teacher network to achieve high accuracy.
problem Machine learning-based sleep apnea detection with limited access to training data.
method Interpretation-based indirect knowledge transfer using activation maximization and synthetic datasets.
result The student classifier achieves 97.8% accuracy on MNIST and 86.1-89.5% on Apnea-ECG dataset.
Bayesian model identifies skill difficulties and student subgroups in engineering education.
problem Identifying and supporting diverse student needs in entry-level university engineering modules.
method Hierarchical Bayesian modeling of student response data.
result Clear patterns of skill mastery and distinct student subgroups identified.
This paper proposes a method to embed teacher knowledge into a student network without increasing parameters.
problem The need for portable neural networks on mobile devices with limited resources.
method Feature embedding approach to distill knowledge from a teacher network to a student network without introducing new parameters.
result The proposed method maintains the performance of the teacher network while significantly reducing computational and storage complexity.
MUTLA dataset analyzes multimodal data for teaching and learning analytics.
problem Lack of a comprehensive multimodal dataset for teaching and learning analytics.
method Presented a large-scale MUTLA dataset with synchronized multimodal data from SAIL.
result Provides insights for predicting student engagement and improving adaptive learning.
Real-time student performance prediction without historical data.
problem Predict student performance in real-time with limited historical data.
method Domain adaptation framework using GritNet architecture with unsupervised transfer learning.
result GritNet generalizes well across different courses and enhances predictions in the early stages.
A new distillation method transfers channel information from teacher to student.
problem Transfer knowledge from teacher to student with fewer parameters and calculations.
method Channel Distillation (CD) and Guided Knowledge Distillation (GKD) with loss decay.
result Achieved 27.68% top-1 error on ImageNet with ResNet18, outperforming state-of-the-art methods.
A method distills GANs for mobile devices, reducing computation and storage.
problem Heavy computation and storage cost of GANs on mobile devices.
method Knowledge distillation to train a smaller generator with inherited information from a larger teacher generator, including a discriminator.
result Portable GAN models with strong performance achieved.
This paper applies machine learning techniques to student modeling. It presents a method for discovering high-level student behaviors from a very large set of low-level traces corresponding to problem-solving actions in a learning environment. Basic actions are encoded into sets of domain-dependent attribute-value patt…