Precise Prosody: Modernizing Greek Emotional Speech Synthesis via Feature Selection

Feature Selection for Improved Phone Duration Modeling of Greek Emotional Speech

2010-01-01
Alexandros Lazaridis, Todor Ganchev, Iosif Mporas, Theodoros Kostoulas, Nikos Fakotakis
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the impact of two feature selection techniques, Relief and Correlation-based Feature Selection (CFS), on phone duration modeling for Greek emotional speech synthesis. By applying ten machine learning algorithms across five emotional categories, the authors demonstrate that reducing high-dimensional feature sets (93 initial features) significantly enhances the accuracy of segmental duration prediction.

TL;DR

Achieving natural-sounding synthetic speech requires precise "phone duration modeling"—knowing exactly how long a sound should last to convey emotions like joy or anger. This study proves that by stripping away irrelevant linguistic noise using Relief and CFS feature selection, we can significantly boost the accuracy of Greek emotional speech models across ten different machine learning architectures, bringing us closer to the "human-like" 20ms error threshold.

The Challenge: The "Noise" in Emotional Prosody

In the world of Text-to-Speech (TTS), prosody is the soul of the voice. It conveys intent and emotion through rhythm and duration. However, modeling this is notoriously difficult because speech features are highly dimensional. When we try to predict how long a "vowel" lasts in a state of Anger vs. Sadness, we often overwhelm models with 90+ variables ranging from syllable position to accentual patterns.

The problem? Many of these features are redundant. This redundancy doesn't just slow down the system; it actively confuses the model, leading to higher Root Mean Square Error (RMSE) and robotic, "flat" emotional delivery.

Methodology: Pruning the Feature Forest

The researchers focused on a Modern Greek emotional database featuring a professional actress expressing five states: Anger, Fear, Joy, Neutral, and Sadness.

1. The Selection Toolkit

They compared two sophisticated algorithms to sift through 93 features:

  • Relief: A statistical approach that assigns weights to features based on how well they distinguish between similar instances, aiming for a "maximal margin."
  • CFS (Correlation-based Feature Selection): This uses a Genetic Algorithm to find the "sweet spot"—a subset of features that are highly correlated with the target (duration) but have low correlation with each other (minimizing redundancy).

2. The Model Battery

They didn't just test one model; they tested ten, including M5p Model Trees, Linear Regression, and Meta-learning (Bagging) algorithms. This provided a comprehensive look at whether feature selection helps everyone or just specific types of learners.

Table 1: RMSE Results Across Emotions

Insights from the Results

The experimental data (as seen in the tables above) revealed several critical technical insights:

  • Consistency is Key: In almost every emotional category (except Joy), using a feature selection algorithm resulted in the lowest error rates.
  • The M5p Advantage: The M5p tree was the "MVP." By using a greedy algorithm to build model trees with linear regression functions at the leaves, it managed to handle the non-linear nature of emotional speech better than standard Linear Regression.
  • Emotional Sensitivity: Different emotions preferred different selection methods. Relief worked best for "Fear" and "Neutral," while CFS was more effective for "Anger" and "Sadness." This suggests that the phonetic "markers" of fear might be more margin-distinct, while anger features are more structurally correlated.

Table 2: Correlation Coefficient (CC) Comparison

Deep Insight: Why Why This Matters for TTS

The industry standard for "high quality" is an error under 20ms. Without feature selection, many models hovered around 22-30ms. By introducing Relief and CFS, the researchers pushed models like the AR.M5pR and M5p into the sub-20ms range (e.g., 19ms for Joy and Fear).

This shift isn't just a statistical win; it's a perceptual one. Reducing duration error directly correlates with a reduction in the "uncanny valley" effect of synthetic emotional speech.

Conclusion & Future Outlook

This work confirms that Feature Selection is a vital preprocessing step in the prosody pipeline. While modern TTS has shifted toward neural architectures (Tacotron, FastSpeech), the underlying principle remains: the quality of prosodic output is fundamentally limited by the relevance of the input features.

Future research could investigate whether these selected feature subsets remain consistent across different languages or if "emotional feature importance" is culturally and linguistically specific to Greek.

Find Similar Papers

Try Our Examples

  • Search for recent studies on phone duration modeling in emotional speech synthesis that utilize deep learning architectures like LSTMs or Transformers to compare with traditional decision tree methods.
  • Which paper first established the 20ms RMSE threshold as the benchmark for perceptual naturalness in speech synthesis, and how has this threshold evolved with neural TTS?
  • Investigate how the Relief and CFS feature selection algorithms have been adapted or replaced by attention-based feature importance techniques in modern prosody modeling.
Contents
Precise Prosody: Modernizing Greek Emotional Speech Synthesis via Feature Selection
1. TL;DR
2. The Challenge: The "Noise" in Emotional Prosody
3. Methodology: Pruning the Feature Forest
3.1. 1. The Selection Toolkit
3.2. 2. The Model Battery
4. Insights from the Results
5. Deep Insight: Why Why This Matters for TTS
6. Conclusion & Future Outlook