Beyond the Report Card: Using GBM to Predict Academic Failure in Brazilian Public Schools

Educational Data Mining: Discovery Standards of Academic Performance by Students in Public High Schools in the Federal District of Brazil

2017-01-01
Eduardo Fernandes, Rommel N. Carvalho, Maristela Holanda, Gustavo Cordeiro Galvão Van Erven
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an Educational Data Mining (EDM) framework to predict student failure in Brazilian public high schools using the CRISP-DM methodology. By employing the Gradient Boosting Machine (GBM) algorithm on the iEducar dataset, the researchers developed models that identify at-risk students even before bimonthly grades are available.

Executive Summary

TL;DR: Researchers in Brazil have moved beyond simple averages to build a predictive "early warning system" for high school students. By applying the Gradient Boosting Machine (GBM) algorithm within the CRISP-DM framework, they successfully predicted student failure rates with high accuracy—identifying that a student's neighborhood and school are often more telling than their first set of grades.

Academic Context: This work sits firmly in the field of Educational Data Mining (EDM). Rather than proving a new theorem, it demonstrates a high-impact application of SOTA ensemble learning (GBM) on massive public sector datasets (iEducar) to solve a critical social problem: student retention.

Problem & Motivation: The Limitation of Descriptive Statistics

For decades, school administrators have relied on Descriptive Statistics—means, medians, and standard deviations. While useful for year-end reports, these metrics are reactive. They tell you how many students failed after the year is over.

The authors argue that waiting for the first bimonthly grades is often too late for effective intervention. The challenge was: Can we predict who will fail on day one of the school year? This requires shifting from simple data summary to complex pattern discovery, finding the latent variables that correlate with failure before the first exam is even taken.

Methodology: CRISP-DM and the Power of Boosting

The project followed the CRISP-DM (Cross-Industry Standard Process for Data Mining), a robust six-phase cyclical process.

The Engine: Gradient Boosting Machine (GBM)

The choice of GBM is strategic. Unlike a single Decision Tree which might overfit or struggle with weak patterns, GBM builds an "ensemble" of trees sequentially. Each new tree attempts to correct the errors of the previous ones.

  • The Input: 17 variables including residency, age, school name, sex, and grant status.
  • The Process: The team built two models. Model 1 used only data available at registration. Model 2 incorporated bimonthly grades and absences.

CRISP-DM Methodology Figure 1: The CRISP-DM cycle used to structure the research.

Experiments & Results: The "Neighborhood" Insight

The model performance was evaluated using ROC Curves (Receiver Operating Characteristic) and Confusion Matrices.

  • Early Year Model: Achieved a ROC of 0.908 on validation data.
  • Mid-Year Model: Achieved a ROC of 0.913 (a slight performance boost as grades were added).

The most striking discovery was the Variable Importance ranking. In the early-year model, the student's neighborhood and the specific school they attended were the most significant predictors. Even when grades were introduced in the second model, "grades" only ranked third in importance.

Variable Importance - Early Year Figure 2: Variable Importance at the beginning of the year, showing the dominance of geographic features over individual demographics.

Model Performance Table Table 1: Confusion matrix showing the model's high sensitivity in identifying potential failures.

Critical Analysis & Conclusion

Takeaway

This research proves that academic failure is not solely an individual performance issue; it is a contextual one. By identifying students based on their environment (neighborhood) and school infrastructure before they even receive their first grade, the State Department of Education can allocate social workers and remedial teachers to specific "at-risk" clusters.

Limitations & Future Work

The study currently lacks the Deployment phase of CRISP-DM—putting the model into a live dashboard for teachers. Future iterations should also consider "Time-to-Failure" analysis and perhaps incorporate Inductive Bias from pedagogical experts to refine the feature engineering process.

By moving from "what happened" to "what will happen," Educational Data Mining serves as a bridge between data science and social equity.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize socio-geographic and demographic metadata to predict high school dropout rates in developing countries.
  • Which paper first established the Gradient Boosting Machine (GBM) algorithm, and how have its implementations in H2O contributed to scalability in large educational datasets?
  • Explore research that applies the CRISP-DM methodology to real-time educational dashboards for teacher-led interventions.
Contents
Beyond the Report Card: Using GBM to Predict Academic Failure in Brazilian Public Schools
1. Executive Summary
2. Problem & Motivation: The Limitation of Descriptive Statistics
3. Methodology: CRISP-DM and the Power of Boosting
3.1. The Engine: Gradient Boosting Machine (GBM)
4. Experiments & Results: The "Neighborhood" Insight
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work