Predicting Engineering Attrition: High-Accuracy Dropout Analysis at UnB Brazil

Educational Data Mining: Analysis of Drop out of Engineering Majors at the UnB - Brazil

2019-12-01
Rodrigo da Fonseca Silveira, Maristela Holanda, Márcio de Carvalho Victorino, Marcelo Ladeira
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the factors behind student attrition in engineering majors at the University of Brasilia (UnB) using Educational Data Mining (EDM). By comparing GLM, GBM, and Random Forest models, the authors identified that Generalized Linear Models (GLM) achieved an accuracy of 86.56% in predicting student dropout.

TL;DR

Researchers at the University of Brasilia (UnB) utilized ten years of student data (2009–2019) to build a predictive model for engineering student attrition. By integrating socio-economic factors like the geographical distance to campus with academic performance, the study deployed a Generalized Linear Model (GLM) that predicts dropout with 86.56% accuracy, forecasting a concerning 57% attrition rate for currently active students.

Contextual Positioning

This work sits at the intersection of Educational Data Mining (EDM) and institutional policy-making. While many EDM papers focus on "black-box" deep learning for performance prediction, this study prioritizes interpretability and novel features (like geocoding zip codes) to provide actionable insights for university administrators.

The Hidden Drivers of Attrition

The "why" behind student dropout is rarely just about bad grades. The authors identified a significant gap in prior literature regarding:

  • Geographical Friction: Does a 40km commute impact retention?
  • Administrative Red Flags: Do multiple requests for a "leave of absence" (qtLeaveAbsenceMajor) signal a higher risk than a single failed exam?
  • Demographic Vulnerability: The unique challenges faced by international and naturalized students.

Methodology: Beyond Standard Performance Metrics

The researchers followed the CRISP-DM (Cross Industry Process Model for Data Mining) framework. They processed data from 5,289 former students and 3,071 active students.

A standout feature of the methodology was the use of GEOCODE software to convert ZIP codes into latitude and longitude, allowing for the calculation of Euclidean distance between a student's home and the Darcy Ribeiro campus.

Model Architecture and Comparison

The study compared three primary algorithms within the H2O Sparkling Water environment:

  1. Generalized Linear Model (GLM): Selected for its high AUC and interpretability.
  2. Gradient Boosting Machine (GBM).
  3. Random Forest (RF).

Model Performance Comparison Fig 1: AUC Comparison showing GLM as the superior predictor for this dataset.

Key Findings and Interpretations

The GLM coefficients provided a clear map of risk factors. Positive coefficients (POS) indicate a higher likelihood of dropout:

  • Physics 1 & Calculus 1: Failing these introductory courses more than twice is a massive predictor of eventual attrition.
  • International Students: Showed a strong positive correlation with dropout, suggesting a need for specialized integration programs (e.g., Portuguese language support).
  • Age Factor: Older students (incomingAge) tend to drop out more frequently, likely due to the dual pressure of work and family life.

GLM Coefficients and Risk Factors Fig 2: Coefficient analysis identifying factors like "foreingStudent" and "rangePhysics1" as high-risk indicators.

Experimental Results

With a final accuracy of 86.56% on the test set, the model proved robust. The confusion matrix revealed that the model is particularly effective at identifying students who will graduate (0.129 error rate for graduates), while being slightly more conservative with dropout predictions.

Critical Insight & Future Outlook

The most striking takeaway is the 57% predicted dropout rate for the current cohort. This is a call to action. The study suggests that "rare" variables—specifically those related to a student's life outside the classroom (distance, housing aid)—are just as critical as academic transcripts.

Limitations: The study currently only uses ZIP codes within the Federal District. Future iterations could benefit from a wider geographical range and the use of a Hadoop cluster to handle larger datasets via Sparkling Water to further refine these predictive "early warning systems."

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize geoprocessing and spatial data to predict university student dropout rates in urban environments.
  • Which original papers established the use of the CRISP-DM methodology specifically for Educational Data Mining (EDM) in higher education?
  • Examine how H2O's Sparkling Water or similar distributed machine learning frameworks are being used to scale dropout prediction models in large-scale public universities.
Contents
Predicting Engineering Attrition: High-Accuracy Dropout Analysis at UnB Brazil
1. TL;DR
2. Contextual Positioning
3. The Hidden Drivers of Attrition
4. Methodology: Beyond Standard Performance Metrics
4.1. Model Architecture and Comparison
5. Key Findings and Interpretations
5.1. Experimental Results
6. Critical Insight & Future Outlook