Beyond the Questionnaire: Predicting Vocational Identity via Socio-Demographic ML

Predicting Vocational Personality Type from Socio-demographic Features Using Machine Learning Methods

2020-10-27
Evgenia Bogacheva, Filipp Tatarenko, Ivan Smetannikov
Summary
Problem
Method
Results
Takeaways
Abstract

This study utilizes supervised machine learning to predict RIASEC vocational personality types (Realistic, Investigative, Artistic, Social, Enterprising, Conventional) using only socio-demographic features. By leveraging a large-scale dataset (n=112,130), the researchers compared multiple architectures, finding that multi-label classification models focusing on label correlations achieved the highest predictive accuracy.

TL;DR

Can your age, education, and geography predict your career path better than a 100-question test? This paper demonstrates that supervised machine learning, specifically multi-label classification that accounts for label correlations, can effectively map socio-demographic features to RIASEC vocational personality types. Using a massive dataset of 112,130 respondents, the researchers achieved significant predictive accuracy (C-Index 13.95 vs 9.0 baseline), proving that our demographic background carries a surprisingly heavy "signature" of our professional interests.

The Shift from Explanation to Prediction

For decades, vocational psychology followed a standard routine: fill out a long survey, get a Holland Code (RIASEC), and find a matching job. As the authors note, psychology has historically prioritized explaining why someone fits a role. However, in the era of Big Data, the focus is shifting toward prediction.

The problem is efficiency. High-fidelity psychological labeling is "expensive" in terms of user time. If we can predict these labels using socio-demographic data—information already present in most social media profiles—we unlock massive potential for automated career counseling and targeted social network analysis.

Methodology: Exploiting the "Circumplex" Structure

The RIASEC model isn't just a list; it’s a circumplex. Scales like "Social" and "Enterprising" are naturally more correlated than "Social" and "Realistic." The researchers hypothesized that models ignoring these relationships would underperform.

They tested four primary approaches:

  1. Independent Regression: Predicting each of the 6 scales separately.
  2. Regression Chains: Predicting scales in a sequence where each subsequent model "sees" the predictions of the previous ones.
  3. Three-Letter Code Classification: Treating the top 3 interests as a single categorical string.
  4. Inferring Label Relations: Using graph-based clustering to find dependencies in the label space.

Actual Pearson Correlation Matrix The figure above illustrates the actual correlations found in the dataset, confirming the inter-dependencies between scales.

Key Insights: What Drives Interests?

By using Gradient Boosting Regressors, the team performed feature importance analysis, revealing a fascinating divide:

  • The Gender Factor: For the Realistic (R) scale (hands-on, mechanical work), gender was the overwhelming predictor.
  • The Geo-Cultural Factor: For the Enterprising (E) scale (leadership, business), individual demographics mattered less than geography, GDP of the home country, and religion.

Feature Importance for R and E scales Table 3: Principal Component Analysis showing how geographical, cultural, and economic dimensions represent the majority of feature variance.

Results and Performance

The breakthrough came with Label Relation Inference. By constructing a label co-occurrence graph (NetworkXLabelGraphClusterer), the model learned the "proximity" of professional interests.

  • Baseline (Dummy): C-Index 9.0
  • Independent Regression: C-Index 10.95
  • Label Relation Inference: C-Index 13.95

The distribution of results showed that the model was notably successful at predicting the exact three-letter Holland code (C-Index 18) far more frequently than chance.

C-Index Distribution Comparison The right-shifted distribution in Figure 6 confirms the model's superior predictive power over the normal distribution of the dummy classifier.

Critical Analysis & Future Outlook

While the results are impressive, the study admits a significant limitation in its metric: the C-Index. This metric penalizes "swapped" letters heavily (e.g., predicting RSA instead of SRA), even though such profiles are practically identical in a counseling context.

The Takeaway: This research moves us closer to a "Career Robot" world. By integrating these ML models into social platforms, we can provide personalized career guidance to users without ever asking them to take a test. For researchers, it highlights that the structure of the output space (label correlations) is just as important as the input features when modeling human psychology.

Find Similar Papers

Try Our Examples

  • Search for recent studies that use Learning to Rank (LTR) algorithms to predict persona-based interest profiles like RIASEC.
  • Which original papers established the "C-Index" and "Hexagonal Congruence Index" for measuring psychological profile similarity, and how have they been adapted for machine learning evaluation?
  • Are there any published works that apply State Space Models or Graph Neural Networks to model the circumplex structure of human personality traits?
Contents
Beyond the Questionnaire: Predicting Vocational Identity via Socio-Demographic ML
1. TL;DR
2. The Shift from Explanation to Prediction
3. Methodology: Exploiting the "Circumplex" Structure
4. Key Insights: What Drives Interests?
5. Results and Performance
6. Critical Analysis & Future Outlook