Decoding Gender in Vietnamese Names: A Deep Learning Benchmark
Gender Prediction Based on Vietnamese Names with Machine Learning Techniques
This paper introduces UIT-ViNames, a novel dataset of over 26,000 annotated Vietnamese full names for gender prediction. The study benchmarks six traditional machine learning algorithms and a Long Short-Term Memory (LSTM) model with fastText embeddings, achieving a SOTA F1-score of 96%.
Executive Summary
TL;DR: Researchers from the University of Information Technology (VNU-HCM) have developed UIT-ViNames, a massive dataset of 26,850 Vietnamese names, and proved that LSTM models combined with fastText embeddings can predict biological gender with an impressive 96% F1-score. The study uniquely highlights that in the Vietnamese context, the "Middle Name" is the secret sauce for high-accuracy classification.
Positioning: This work fills a significant gap in regional NLP. While most global systems struggle with non-Western name structures, this paper provides both the data and the optimized architectural blueprint specifically for the Vietnamese linguistic landscape.
The "Nguyen" Problem: Why Surnames Don't Matter
In many Western cultures, a surname can sometimes hint at heritage, but in Vietnam, the surname distribution is extremely skewed. Approximately 40% of the population shares the surname "Nguyá»…n".
The authors' initial analysis (Figure 1) confirms a critical intuition: the distribution of surnames between males and females is virtually identical. Using a surname to predict gender in Vietnam is mathematically equivalent to a random guess.
Figure 1: Comparison showing that common Vietnamese surnames (Nguyễn, Trần, Lê) provide zero discriminatory power for gender.
Methodology: Beyond Simple Dictionaries
The study compared traditional "Bag-of-Words" approaches with sequential Deep Learning.
1. Traditional Baselines
The team tested SVM, Multinomial Naive Bayes, Bernoulli NB, Logistic Regression, and Random Forest. Using TF-IDF and Count Vectorization, these models performed admirably (around 94-95% accuracy), but struggled with "unisex" names.
2. The Winning Combo: LSTM + fastText
The core of the successful approach was a Long Short-Term Memory (LSTM) network. Unlike traditional models, LSTMs can maintain the "memory" of name order. By using fastText embeddings (300-dimension), the model captures the morphological relationships between names even when they haven't been seen in the training set.
Figure 3: The distribution of common first names highlights gender-specific clusters (e.g., "Thị" for females, "Văn" for males).
Experimental Insights: The Power of the Middle Name
The most technical revelation of the paper comes from the Ablation Study. The researchers systematically stripped away parts of the name to see what happened to the accuracy:
| Name Combination | SVM (Avg F1) | LSTM (Avg F1) |
|---|---|---|
| Family Name Only | 39.60% | 38.23% |
| First Name Only | 87.04% | 80.02% |
| Middle + First Name | 95.28% | 95.89% |
Key Insight: Standalone first names (like "Anh" or "Tú") are often gender-neutral. Adding the Middle Name acts as a disambiguator. For example, "Tuấn Anh" is almost certainly male, while "Tú Anh" is frequently female.
Critical Analysis & Conclusion
The "Unseen" Challenge
Despite the high accuracy, the model still fails on rare or modern naming conventions. For instance, the name "Lâm Anh" (typically female) was misclassified when assigned to a male in the dataset. These "edge cases" represent the cultural shift in Vietnam toward more unique, non-traditional names.
Takeaways
- Architecture Matters: For short-text sequences like names, LSTMs provide the necessary sequential context that Naive Bayes lacks.
- Data Sparsity: Surnames in Vietnamese are "noise"—removing them simplifies the model and potentially improves generalization.
- Future Path: The authors suggest moving toward Transformer-based models (BERT) to see if bidirectional context can push that 96% closer to 100%.
This research provides a robust API-ready solution for automated form-filling, coreference resolution, and demographic analysis in the Vietnamese digital ecosystem.
