VGG-16 vs. MobileNet: Benchmarking Gender Classification for Asian Demographics
Gender Classification Based on Asian Faces using Deep Learning
This paper explores gender classification specifically targeting Asian faces by benchmarking three Deep Learning architectures: VGG-16, ResNet-50, and MobileNet. Leveraging transfer learning with ImageNet pre-trained weights, the study identifies VGG-16 as the most effective model for this demographic, achieving a 100% training accuracy and an 88% recognition rate on real-world test images.
TL;DR
This research addresses the "representation gap" in facial recognition by focusing on Asian demographics. By benchmarking three industry-standard models—VGG-16, ResNet-50, and MobileNet—the authors demonstrate that while lightweight models are efficient, VGG-16 remains the superior choice for accuracy, achieving an 88% recognition rate on a custom Asian face dataset compared to a disappointing 49% from MobileNet.
Context & Motivation
Despite the explosion of AI-driven surveillance and human-computer interaction (HCI) tools, gender classification accuracy remains inconsistent across different ethnicities. Most standard models are trained on datasets with a heavy Caucasian bias (like LFW or Adience).
The authors identify two major hurdles:
- Lack of Diversity: Existing SOTA models often fail on "real-world" Asian faces due to data distribution shifts.
- Localized Data Scarcity: There is a lack of publicly available, high-quality facial databases specifically featuring Malaysian and wider Asian demographics.
Methodology: The Transfer Learning Approach
The study utilizes Transfer Learning, taking models pre-trained on the massive ImageNet dataset and fine-tuning them on a specialized niche.
The Pipeline
- Pre-processing: Faces are detected and extracted using the Haar Cascade algorithm via OpenCV.
- Architecture Selection:
- VGG-16: Known for its simplicity and depth (138 million parameters).
- ResNet-50: Utilizes "skip connections" to mitigate vanishing gradients.
- MobileNet: Designed for mobile efficiency using depth-wise separable convolutions.
- Training: Fixed at 100 epochs with a learning rate of 0.001 using the SGD optimizer.
Fig 1: The proposed classification workflow from image input to probability output.
Experiments & Performance Analysis
The authors built the U10 Face Database, consisting of 1,000 images (50/50 male-female split). The results highlight a stark contrast between model complexity and performance on this specific task.
Training Convergence
All models showed high capability on the training set (all >99%), but the real test was the Recognition Rate on unseen faces.
| Model | Training Accuracy | Actual Recognition Rate | Parameters |
|---|---|---|---|
| VGG-16 | 100% | 88% | 138.3M |
| ResNet-50 | 99.9% | 85% | 25.6M |
| MobileNet | 99.8% | 49% | 4.2M |
The "MobileNet Failure" Insight
The most striking result is MobileNet's 49% accuracy—essentially no better than a coin flip for a binary class. This suggests that the depth-wise separable convolutions, while excellent for saving power, may discard the fine-grained spatial textures required to distinguish gender in Asian faces when the dataset size is limited.
Fig 2: Training loss visualization showing rapid convergence across all models.
Critical Analysis & Takeaways
The paper confirms that for high-stakes biometric identification, bigger is often better when it comes to feature extraction power. VGG-16's dense parameterization allows it to capture subtle facial markers that ResNet and MobileNet missed.
Limitations:
- Dataset Size: 1,000 images is relatively small for deep learning; this likely contributed to the high training accuracy (overfitting) and the struggle of the lightweight MobileNet.
- Demographic Granularity: While focusing on "Asian faces," further sub-categorization (East Asian vs. South Asian) could yield even more specific insights into model bias.
Future Work: The authors suggest exploring InceptionResNetV2 and DenseNet to see if feature concatenation or multi-scale processing can push the 88% recognition rate closer to 95%+.
Conclusion
This study serves as a vital benchmark for localized AI applications. It proves that while we strive for "mobile-first" AI, we cannot sacrifice the depth required to handle demographic diversity. For developers building HCI systems in Asia, VGG-based backbones remain the gold standard for reliable gender classification.
