Debunking the Employment Myth: Data Mining Academic Success in Computer Science
Academic performance of university students and its relation with employment
This paper applies Educational Data Mining (EDM) to analyze student academic performance at the University of La Plata. Using a Wrapper-based feature selection method and Support Vector Machines (SVM), it identifies key predictors of student regularity and develops descriptive models via C4.5 and PART algorithms.
TL;DR
Is a job the enemy of the diploma? A deep dive into a decade of student data from the University of La Plata (UNLP) suggests otherwise. By employing a Wrapper-based feature selection and Machine Learning classifiers, researchers found that actual employment does not decrease academic performance. Instead, the true predictors of falling into "non-regular" status are age, the gap between high school and university, and the psychological state of seeking work.
The "Curse of Choice" in Educational Data
Academic institutions are drowning in data but starving for knowledge. While systems like SIU-Guaranà track everything from socio-economics to language levels, the high dimensionality of this data often creates "noise."
The researchers identified two major hurdles:
- Complexity: High-dimensional data produces massive, uninterpretable decision trees.
- Class Imbalance: With 71% of students remaining "regular," standard classifiers often ignore the minority "at-risk" group.
Methodology: Trimming the Fat
The authors leveraged a KDD (Knowledge Discovery in Databases) workflow. To move from raw data to a predictive model, they transformed sparse variables (like specific high school names) into categorical ones (public vs. private) and introduced the "Avance" (Progress) metric:
Where represents approved exams and the total career requirements.
The core of the methodology lies in the Chi2 Feature Selection. By calculating the relationship between features and the target class, they reduced the dataset to a "Core 4":
- Seeking Work (Busca Trabajo)
- Age (Edad)
- Time to Enter University (Tiempo en ingresar)
- Academic Pace (Ritmo)
Figure 1: Missing data patterns (dark areas) helped prune the initial search space.
Counter-Intuitive Findings: Work vs. Performance
The most striking result from the C4.5 and PART models (reaching ~80% accuracy) is what is missing from the predictive list.
Contrary to popular belief, the number of hours worked and employment status did not survive the feature selection process. This implies a "Self-Regulation Insight": students who work likely manage their time more effectively or adjust their course load to match their capacity.
However, the intent to work (seeking a job) coupled with a long gap ( >2.5 years) between high school and college significantly increased the risk of losing "regular" status.
Figure 2: Model comparison showing that using only selected features (top) maintained similar accuracy to using all features while improving the precision of identifying "At-Risk" (Non-Regular) students.
Critical Analysis & Conclusion
Takeaway
The paper shifts the focus from socio-economic status to educational trajectory. The "at-risk" profile isn't just a "working student"; it’s the student who is older, delayed their entry into the faculty, and is currently preoccupied with the job hunt.
Limitations
- Domain Specificity: The study is localized to Computer Science. CS students often find work easier than those in humanities, which might skew the "work doesn't hurt" conclusion.
- Data Latency: The questionnaire is filled out at registration; changes in employment status during the degree might not be captured dynamically.
Future Outlook
This work sets a precedent for Predictive Tutoring. By identifying these 4 key variables at registration, universities can flag at-risk students on "Day 1" and provide targeted counseling before the first exam is even failed.
