Predicting Academic Fate: A Systematic Mapping of Educational Problem Forecasting

Approaches to Predicting Educational Problems: A Systematic Mapping

2020-10-21
Paulo Silva, Fernando da Fonseca de Souza, Roberta Fagundes
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a Systematic Mapping Literature (SML) study investigating methodologies for predicting educational problems like dropout and low academic performance. By synthesizing research from 2010 to 2019, it identifies Educational Data Mining (EDM), Machine Learning (ML), and Learning Analytics (LA) as the primary technical frameworks achieving state-of-the-art predictive results.

TL;DR

In an era of educational Big Data, predicting whether a student will fail or drop out has moved from guesswork to a rigorous computational science. This paper maps a decade of progress (2010-2019), revealing that while Classification and Regression models are becoming highly accurate, the field still struggles to explain why students fail in a way that aligns with educational theory.

The Growing Crisis in Education

Educational institutions globally are battling "alarming" rates of dropout and learning stagnation. The shift to digital environments (LMS/VLE) has created a goldmine of data, but identifying at-risk students remains a "wicked problem." The authors argue that while we can predict who might fail, we are still far from understanding the structural and behavioral causes behind these outcomes in a standardized way.

Methodology: Mapping the Landscape

The authors conducted a Systematic Mapping Literature (SML) review, filtering hundreds of papers down to 50 core studies that utilized Educational Data Mining (EDM), Learning Analytics (LA), and Machine Learning (ML).

1. The Taxonomy of Predictions

The research categorized studies into three distinct "facets":

  • Performance Prediction: Focusing on grades and engagement.
  • Dropout (Evasion) Prediction: Early identification of students likely to leave.
  • Behavioral Prediction: Analyzing how students interact with virtual systems (primarily for Distance Education).

2. The Data Goldmine

Where does this knowledge come from? The mapping shows a healthy mix of data sources: Data Sources Distribution

  • LMS/VLE Interactions (22.4%): Clicks, forum posts, and time-on-task.
  • Demographics (20.9%): Age, gender, and social class.
  • Academic History (19.4%): Previous grades and failed credits.

Core Techniques: The Engine of Prediction

The paper highlights a heavy reliance on supervised learning. Classification was used in the majority of papers to segment students into "risk profiles," while Regression was used to isolate specific variables that impact performance.

Techniques Comparison

A notable trend is the use of Ensemble Methods. By combining different algorithms (like Random Forest with Linear Regression), researchers are effectively reducing error rates and increasing the robustness of early warning systems.

Critical Analysis: What’s Missing?

Despite the technical wins, the Senior Editor’s perspective notes three critical gaps:

  1. The K-12 Gap: Most research is concentrated on Higher Education. Predictive models for early childhood and secondary education are significantly under-represented.
  2. Theoretical Disconnect: There is a "black box" problem. Algorithms identify correlations (e.g., "students who log in late fail"), but they don't necessarily prove causation or tie back to socio-psychological theories of learning.
  3. Standardization: There is no "one-size-fits-all" model. Every institution uses different collection methods, making it hard to generalize a high-performing model from one university to another.

Conclusion & Future Outlook

This mapping confirms that Educational Data Mining is no longer a niche interest—it is a cornerstone of modern institutional strategy. However, the next frontier isn't just "better accuracy." It is Interpretability. The industry needs dashboards that don't just show a "red flag" for a student, but explain the socioeconomic or behavioral factors behind it, allowing educators to intervene with precision.

Takeaway for Researchers: If you are building the next SOTA model, stop looking for more data—start looking for ways to align your features with established educational frameworks to provide "actionable" insights.

Find Similar Papers

Try Our Examples

  • Search for recent papers that integrate theoretical educational models like Tinto's Student Integration Model with Machine Learning for dropout prediction.
  • What are the current SOTA methods for "Early Warning Systems" in K-12 education compared to the Higher Education focus found in Silva et al. (2020)?
  • Examine how Explainable AI (XAI) is being applied to Educational Data Mining to solve the "cause-and-effect" gap identified in this mapping study.
Contents
Predicting Academic Fate: A Systematic Mapping of Educational Problem Forecasting
1. TL;DR
2. The Growing Crisis in Education
3. Methodology: Mapping the Landscape
3.1. 1. The Taxonomy of Predictions
3.2. 2. The Data Goldmine
4. Core Techniques: The Engine of Prediction
5. Critical Analysis: What’s Missing?
6. Conclusion & Future Outlook