Decoding Employability: A Data Mining Approach to Graduate Success

A Classification-Based Graduates Employability Model for Tracer Study by MOHE

2011-01-01
Myzatul Akmam Sapaat, Aida Mustapha, Johanna Ahmad, Khadijah Chamili, Rahamirzam Muhamad
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a systematic data mining approach to construct a Graduates Employability Model using classification techniques. Focused on the Malaysian Ministry of Higher Education (MOHE) Tracer Study data, the authors benchmarked Bayes-based and tree-based algorithms, identifying J48Graft as the SOTA classifier with 92.3% accuracy for predicting post-graduation employment status.

TL;DR

Higher education institutions produce thousands of graduates annually, but predicting their market readiness remains a complex challenge. This paper leverages the Malaysian MOHE Tracer Study (12,830 instances) to build a predictive model for employability. By comparing Bayes and Tree-based algorithms, the research finds that the J48Graft algorithm achieves a superior 92.3% accuracy, identifying industry sectors and job status as more critical than traditional academic metrics.

The "Why": Moving Beyond Qualitative Surveys

For years, employability research has been the domain of social scientists using manual interviews and small-scale surveys. However, as the volume of graduates grows, these methods fail to capture the "hidden patterns" within national databases. The authors identify a critical gap: despite having access to the massive MOHE Tracer Study database, the industry lacked a robust, automated classification model to predict whether a graduate would end up employed, unemployed, or in an "undetermined" state.

Methodology: Bayes vs. Trees

The study treats employability as a supervised classification task. The researchers performed rigorous preprocessing, including:

  • Data Discretization: Converting continuous values like CGPA and Age into categorical intervals to optimize classifier performance.
  • Feature Engineering: Using Information Gain to rank 20 different attributes.

The Battle of Algorithms

The core of the paper lies in the head-to-head comparison between two distinct mathematical philosophies:

  1. Bayesian Methods: Statistical models (like Naïve Bayes and AODE) that estimate prior and posterior distributions to assign class probabilities.
  2. Tree-based Methods: Logic-driven models (like J48 and REPTree) that sort instances down branches based on attribute tests.

Model Architecture: Decision Tree Representation

Insights from Information Gain

One of the most striking findings is the ranking of influencing factors. While students often obsess over CGPA, the Information Gain analysis showed that the Job Sector and Job Status are the most powerful discriminators for classification.

Variable Importance through Information Gain Fig 1: A radial display showing the Root Mean Squared Error (RMSE) across algorithms. Tree-based methods (lower RMSE) generally provided better forecasts than Bayes counterparts.

Experimental Results & SOTA Performance

The researchers found that J48Graft (a variant of the C4.5 tree) achieved the highest accuracy. The secret to its success is Grafting: unlike standard pruning which removes branches to simplify the tree, grafting adds nodes using non-local information to identify predictive patterns in regions of the data that are sparsely populated.

AlgorithmAccuracy (%)Kappa Statistic
J48Graft92.30.849
J4892.20.848
AODE (Best Bayes)91.10.827
Naïve Bayesian90.90.825

Critical Analysis & Conclusion

Takeaway

The study proves that tree-based classifiers are more "knowledge-friendly" for educational datasets. They don't just provide a prediction; they provide a readable logic path (the tree structure) that policymakers can use to understand why certain graduates are struggling.

Limitations & Future Work

While the accuracy is high, the authors acknowledge a "confidentiality gap"—10% of attributes were unavailable due to sensitivity issues. Furthermore, the model is a snapshot of the first six months post-graduation.

The next frontier for this research involves Clustering-based Preprocessing and the integration of Alumni Data to create a longitudinal view of employability that evolves as a career progresses.

Professional Perspective

This work represents a solid transition from descriptive statistics to predictive analytics in educational administration. By identifying J48Graft as the optimal classifier, it sets a technical benchmark for future Intelligent Student Information Systems (ISIS) in the ASEAN region.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize advanced ensemble methods like XGBoost or LightGBM on national graduate tracer studies to improve classification accuracy beyond traditional decision trees.
  • Which study first introduced the "grafting" technique for C4.5 decision trees (J48Graft), and how does it specifically mitigate the fragmentation problem in leaf nodes compared to standard pruning?
  • Explore research that integrates Longitudinal Data Analysis with Graduate Employability models to track career progression beyond the initial six-month post-graduation window.
Contents
Decoding Employability: A Data Mining Approach to Graduate Success
1. TL;DR
2. The "Why": Moving Beyond Qualitative Surveys
3. Methodology: Bayes vs. Trees
3.1. The Battle of Algorithms
4. Insights from Information Gain
5. Experimental Results & SOTA Performance
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations & Future Work
6.3. Professional Perspective