Beyond the Report Card: Using GBM to Predict Academic Failure in Brazilian Public Schools
Educational Data Mining: Discovery Standards of Academic Performance by Students in Public High Schools in the Federal District of Brazil
This paper presents an Educational Data Mining (EDM) framework to predict student failure in Brazilian public high schools using the CRISP-DM methodology. By employing the Gradient Boosting Machine (GBM) algorithm on the iEducar dataset, the researchers developed models that identify at-risk students even before bimonthly grades are available.
Executive Summary
TL;DR: Researchers in Brazil have moved beyond simple averages to build a predictive "early warning system" for high school students. By applying the Gradient Boosting Machine (GBM) algorithm within the CRISP-DM framework, they successfully predicted student failure rates with high accuracy—identifying that a student's neighborhood and school are often more telling than their first set of grades.
Academic Context: This work sits firmly in the field of Educational Data Mining (EDM). Rather than proving a new theorem, it demonstrates a high-impact application of SOTA ensemble learning (GBM) on massive public sector datasets (iEducar) to solve a critical social problem: student retention.
Problem & Motivation: The Limitation of Descriptive Statistics
For decades, school administrators have relied on Descriptive Statistics—means, medians, and standard deviations. While useful for year-end reports, these metrics are reactive. They tell you how many students failed after the year is over.
The authors argue that waiting for the first bimonthly grades is often too late for effective intervention. The challenge was: Can we predict who will fail on day one of the school year? This requires shifting from simple data summary to complex pattern discovery, finding the latent variables that correlate with failure before the first exam is even taken.
Methodology: CRISP-DM and the Power of Boosting
The project followed the CRISP-DM (Cross-Industry Standard Process for Data Mining), a robust six-phase cyclical process.
The Engine: Gradient Boosting Machine (GBM)
The choice of GBM is strategic. Unlike a single Decision Tree which might overfit or struggle with weak patterns, GBM builds an "ensemble" of trees sequentially. Each new tree attempts to correct the errors of the previous ones.
- The Input: 17 variables including residency, age, school name, sex, and grant status.
- The Process: The team built two models. Model 1 used only data available at registration. Model 2 incorporated bimonthly grades and absences.
Figure 1: The CRISP-DM cycle used to structure the research.
Experiments & Results: The "Neighborhood" Insight
The model performance was evaluated using ROC Curves (Receiver Operating Characteristic) and Confusion Matrices.
- Early Year Model: Achieved a ROC of 0.908 on validation data.
- Mid-Year Model: Achieved a ROC of 0.913 (a slight performance boost as grades were added).
The most striking discovery was the Variable Importance ranking. In the early-year model, the student's neighborhood and the specific school they attended were the most significant predictors. Even when grades were introduced in the second model, "grades" only ranked third in importance.
Figure 2: Variable Importance at the beginning of the year, showing the dominance of geographic features over individual demographics.
Table 1: Confusion matrix showing the model's high sensitivity in identifying potential failures.
Critical Analysis & Conclusion
Takeaway
This research proves that academic failure is not solely an individual performance issue; it is a contextual one. By identifying students based on their environment (neighborhood) and school infrastructure before they even receive their first grade, the State Department of Education can allocate social workers and remedial teachers to specific "at-risk" clusters.
Limitations & Future Work
The study currently lacks the Deployment phase of CRISP-DM—putting the model into a live dashboard for teachers. Future iterations should also consider "Time-to-Failure" analysis and perhaps incorporate Inductive Bias from pedagogical experts to refine the feature engineering process.
By moving from "what happened" to "what will happen," Educational Data Mining serves as a bridge between data science and social equity.
