Beyond Voice Pitch: Improving Gender Identification with Multi-Task Age Context

Multi-task learning DNN to improve gender identification from speech leveraging age information of the speaker

2020-02-10
Mousmita Sarma, Kandarpa Kumar Sarma, Nagendra Kumar Goel
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a Multi-Task Learning (MTL) Deep Neural Network (DNN) that improves gender identification from speech by using the speaker's age as an auxiliary target. The system employs an end-to-end architecture featuring a raw waveform front-end with 1D-Convolutional layers and LSTMP (Long Short Term Memory with Recurrent Projection) layers for temporal modeling.

TL;DR

Researchers have developed a new Deep Neural Network (DNN) architecture that solves a classic problem in speech analysis: identifying gender across all ages. By teaching a model to recognize both age and gender simultaneously from raw audio waves (avoiding traditional MFCCs), they achieved a massive 29.9% relative error reduction compared to traditional methods on diverse real-world datasets.

Contextualizing the Challenge

Gender identification might seem like a "solved" problem in AI, with many models hitting 98%+ accuracy on adult speech. However, these models often fail when encountering children or seniors. In kids, the physiological differences in vocal tracts haven't fully diverged yet, and in seniors, hormonal and physical changes can blur vocal gender lines.

The authors of this paper realized that Human Psychology rarely views gender in a vacuum—we naturally adjust our expectations based on the speaker's perceived age. They set out to replicate this inductive bias in a machine learning framework.

The "Raw" Methodology

The proposed system departs from the traditional pipeline of "Feature Extraction (MFCC) → Classifier" in two fundamental ways:

1. The Multi-Task Learning (MTL) Paradigm

Instead of just predicting "Male or Female," the network has two "heads."

  • Primary Head: Predicts Gender.
  • Auxiliary Head: Predicts the Age-Gender group (e.g., "Young Male" vs "Senior Female").

This forces the internal "hidden" layers of the AI to learn a representation that understands how age influences voice before it makes a final judgment on gender.

2. Learning from the Source

Most speech AI uses MFCCs—a "hand-crafted" summary of audio. This paper uses the Raw Waveform.

  • Front-End: A 1D-Convolutional Neural Network (CNN) acts as a learnable filter bank, discovering acoustic features directly from the time-domain signal.
  • Temporal Modeling: LSTMP (LSTM with Recurrent Projection) layers are used to track long-term patterns in the speech, outperforming the more common TDNN (Time Delay Neural Network) architectures.

Overall Architecture Figure 1: The Multi-task learning DNN layout showing shared hidden layers and the dual-output head.

Experimental Proof

To test the model, the team combined data from NIST SRE (adults) and the OGI Kids corpus. The results were telling:

ModelWeighted Accuracy (WA)Unweighted Accuracy (UA)
Traditional GMM (MFCC)86.74%86.80%
Single-Task DNN (Raw Wave)87.89%87.77%
Multi-Task DNN (Raw Wave)90.70%90.65%

The "Aha!" moment comes from the Ablation Study. When comparing the Single-Task model (Gender only) to the Multi-Task model (Gender + Age), the error rate dropped by 23.2%. This proves that the age information isn't just "extra data"—it's a critical guide for the network's understanding.

Performance across Age Groups Table: Comparison shows the Multi-Task model generalizing significantly better across children and seniors.

Critical Insight & Future Outlook

This work highlights a shift in Speech AI: moving away from pre-defined mathematical features (like MFCC) toward end-to-end learning. More importantly, it shows that "Para-linguistic" traits (age, gender, emotion) are not independent.

Limitations: While gender identification improved, the authors noted that age classification did not necessarily see a boost from gender data. This suggests that while age is a precursor to understanding gender, the reverse might not be as numerically significant for the model.

Conclusion: For developers of Smart Call Centers or Human-Machine Interaction (HMI) systems, this paper provides a blueprint for building more inclusive AI that doesn't "misgender" a speaker simply because they are a child or an elderly person.

Find Similar Papers

Try Our Examples

  • Find recent papers published after 2020 that use Multi-Task Learning to jointly identify age, gender, and emotion from raw speech signals.
  • Which study first introduced the concept of LSTMP (LSTM with Recurrent Projection) layers in acoustic modeling and how does it compare to standard LSTMs in parameter efficiency?
  • Search for research that applies similar raw waveform convolutional front-ends to cross-lingual speaker recognition tasks to verify domain generalization.
Contents
Beyond Voice Pitch: Improving Gender Identification with Multi-Task Age Context
1. TL;DR
2. Contextualizing the Challenge
3. The "Raw" Methodology
3.1. 1. The Multi-Task Learning (MTL) Paradigm
3.2. 2. Learning from the Source
4. Experimental Proof
5. Critical Insight & Future Outlook