Bridging the Emotional Gap: Scaling Speech Recognition with Twitter Intelligence

Language Model Adaptation for Emotional Speech Recognition using Tweet data

2020-12-07
Kazuya Saeki, M. Katoh, T. Kosaka
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a language model (LM) adaptation method using a large-scale Twitter dataset (25.86M words) to improve Japanese Emotional Speech Recognition. By leveraging the colloquial and emotional nature of tweets, the authors significantly reduce the Word Error Rate (WER) on the JTES and OGVC corpora.

TL;DR

Recognizing emotional speech has long been a "stumbling block" for ASR systems due to the vast difference between formal training data and real-world outbursts. This paper presents a breakthrough by using 25.86 million words of Twitter data to adapt Language Models, cutting Word Error Rates (WER) from 36.11% down to 17.77%.

Background: Why "Normal" ASR Fails at Emotion

Most Japanese ASR systems are trained on the Corpus of Spontaneous Japanese (CSJ)—essentially a collection of academic lectures. While great for formal settings, it is "emotionally sterile." When a speaker gets angry, joyful, or sad, two things happen:

  1. Acoustic Shift: Pitch, intensity, and duration change drastically.
  2. Linguistic Shift: Speakers use "colloquialisms," drop particles (like ga or wo), and emphasize specific Japanese phonemes like the geminate stop /Q/ (baQkari).

Previous studies failed because they tried to adapt models using tiny datasets (under 2,000 sentences). The authors’ core insight is that volume beats precision: a massive amount of unlabeled, informal text (Tweets) is more valuable than a handful of perfectly labeled emotional sentences.

Methodology: The Two-Pronged Adaptation

1. Large-Scale Twitter LM Adaptation

The authors used the Twitter API to collect 51 days of Japanese tweets. After stripping URLs and hashtags, they applied a Bigram Perplexity filter using the CSJ model to ensure the selected sentences were linguistically coherent while retaining their "human" emotional flavor.

The adaptation uses a Mixed N-gram approach, which mathematically balances the baseline (formal) counts with the new (tweet) counts: This allows the model to "remember" standard Japanese grammar while "learning" the emotional shortcuts used on social media.

2. Acoustic Model (AM) Adaptation

The system utilizes a DNN-HMM architecture. To handle the acoustic variance, the authors performed supervised backpropagation on the JTES (Japanese Twitter-based emotional speech) corpus. A key innovation here is Output Probability Compensation, which prevents common states (like silence) from overwhelming the model's predictions during emotional peaks.

System Architecture & Perplexity improvement Table: Test set perplexity significantly improves (from ~900 to ~224) when switching to large-scale tweet adaptation.

Experimental Results: Quantitative Dominance

The results on the JTES corpus were striking across all emotion types (Anger, Joy, Neu, Sad):

MethodAverage WER (%)
Baseline (CSJ)36.11
LM Adaptation (Large-Scale)25.68
Combined AM + LM Adaptation17.77

Accuracy by Emotion Table: Results show that "Sadness" and "Anger" saw the most dramatic improvements, likely due to the higher frequency of colloquial markers in those states.

Qualitative Win: Handling the "Nuance"

The paper highlights specific cases where the model succeeded:

  • Particle Dropping: Correctly identifying "zikan aru toki" instead of the formal "zikan ga aru toki."
  • Emphasis: Capturing the emphatic stop /Q/ in "baQkari."
  • Morphofogical Accuracy: Avoiding errors in auxiliary verbs like "na no ni."

Deep Insight & Conclusion

The most significant takeaway is Versatility. Even when tested on the OGVC (Online Gaming Voice Chat) corpus—a completely different environment from Twitter—the model still outperformed the baseline. This suggests that the "emotional language" of Twitter is a generic proxy for "informal human interaction."

Limitations: The model still struggles with "fillers" (long vowels) and extreme emotional intensity (Intensity Level 3). The authors suggest that future work should focus on Emotion-Dependent AMs and upgrading from N-grams to Transformer-based Neural LMs.

Takeaway for Practitioners: If you are building ASR for real-world interactions, don't just look for speech data. A massive crawl of informal text from the target culture may be the cheapest and most effective "performance booster" for your language model.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Large Language Models (LLMs) like GPT or BERT for emotional and colloquial Japanese text data augmentation in ASR systems.
  • What are the current state-of-the-art methods for "cross-corpus" emotional speech recognition where training and testing data come from different acoustic environments?
  • Explore how End-to-End (E2E) ASR architectures, such as Whisper or Conformer, handle the specific linguistic markers of emotional Japanese (e.g., geminate consonants or particle dropping) compared to DNN-HMM systems.
Contents
Bridging the Emotional Gap: Scaling Speech Recognition with Twitter Intelligence
1. TL;DR
2. Background: Why "Normal" ASR Fails at Emotion
3. Methodology: The Two-Pronged Adaptation
3.1. 1. Large-Scale Twitter LM Adaptation
3.2. 2. Acoustic Model (AM) Adaptation
4. Experimental Results: Quantitative Dominance
4.1. Qualitative Win: Handling the "Nuance"
5. Deep Insight & Conclusion