Decoding the Learner's Path: Predicting Performance via Pedagogical Video Navigation
Predicting learner’s performance through video sequences viewing behavior analysis using educational data-mining
This paper introduces a performance prediction framework using Educational Data Mining (EDM) to analyze learner interactions with pedagogical video sequences. By shifting focus from generic click types to sequence-specific navigation, the authors leverage K-Nearest Neighbors (k-NN) and Multilayer Perceptron (MLP) to achieve an average classification accuracy of 65.07%, effectively predicting 'Pass' or 'Fail' outcomes based on viewing behavior.
TL;DR
With the surge of distance learning post-COVID-19, understanding how students interact with educational videos is critical. This study moves beyond counting clicks to analyzing where (in which pedagogical sequence) those clicks occur. Using k-NN and MLP algorithms, the researchers achieved a 65%+ accuracy rate in predicting whether a student would pass or fail, purely based on their navigation behavior through structured video segments.
Problem & Motivation: The Contextual Gap in Video Analytics
Most existing studies in Educational Data Mining (EDM) focus on quantitative metrics: How many times did the user press "Pause"? How many "Fast Forwards" occurred? While valuable, these metrics ignore the instructional context.
The authors argue that a "Pause" during an introductory segment carries different cognitive weight than a "Pause" or "Backward Jump" during a complex mathematical derivation. Existing methods often miss this nuance, treating the video as a flat timeline rather than a hierarchy of concepts.
Methodology: Mapping Behavior to Content
The core innovation lies in the Pedagogical Sequence Segmentation. Rather than treating the video as a single 15-minute file, it is divided into logical fragments (e.g., Introduction Examples Deep Dive).
1. The Processing Pipeline
The authors followed a traditional EDM workflow but injected sequence-awareness at the preprocessing stage:
- Data Collection: 66 learners, 4 C++ video courses via Moodle.
- Instrumentation: A customized "Vidtrack" plugin recorded not just the click, but the sequence ID where it happened.
- Feature Transformation: Raw clickstreams were converted into "Path Sequences" (e.g., Sequence 1 Jump back to Seq 1 Proceed to Seq 2).
Figure 1: Manual segmentation of video content into pedagogical units.
2. Predictive Modeling
Two primary algorithms were tested:
- K-Nearest Neighbors (k-NN): A "lazy learner" that classifies a student's path based on the most similar paths in the training set.
- Multilayer Perceptron (MLP): A neural network approach using the DeepLearning4j framework to find non-linear relationships between navigation and grades.
Experimental Results: The Predictability of Success
The study proved that viewing behavior is a viable predictor of performance.
| Classifier | Average Accuracy | Best Performance (Video Specific) |
|---|---|---|
| k-NN | 65.07% | 67.27% |
| MLP | 61.13% | 66.67% |
Key Insights from the Data:
- Imbalance in Detection: The models were much better at predicting 'Pass' students (TP rates up to 0.90) than 'Fail' students. This suggests that "successful" viewing patterns are more homogeneous and easier to model than the varied ways in which students struggle.
- Algorithm Efficiency: Despite the complexity of MLP, the simpler k-NN algorithm provided more consistent reliability (higher ROC Area values) across different video modules.
Figure 2: Accuracy variance across different video courses for k-NN and MLP.
Critical Analysis & Conclusion
This work provides a strong foundation for Early Warning Systems (EWS) in e-learning. By identifying "deviation" in viewing behavior—such as excessive jumping back in a specific pedagogical sequence—instructors can provide targeted guidance before the student even takes a quiz.
Limitations & Future Directions
While the logic is sound, the study has two main areas for growth:
- Manual Effort: Currently, pedagogical sequences must be defined manually by teachers. For this to scale to thousands of videos, Automated Video Content Analysis (using AI to detect scene changes and topic shifts) is necessary.
- Dwell Time: The current model focuses on clicks. Integrating "Time Spent" per sequence would likely boost the accuracy significantly, as it distinguishes between a quick skip and a focused re-watch.
The Takeaway: If you want to know if a student is learning, don't just see if they watched the video—look at how they navigated the ideas within it.
