Precise Prosody: Modernizing Greek Emotional Speech Synthesis via Feature Selection
Feature Selection for Improved Phone Duration Modeling of Greek Emotional Speech
This paper investigates the impact of two feature selection techniques, Relief and Correlation-based Feature Selection (CFS), on phone duration modeling for Greek emotional speech synthesis. By applying ten machine learning algorithms across five emotional categories, the authors demonstrate that reducing high-dimensional feature sets (93 initial features) significantly enhances the accuracy of segmental duration prediction.
TL;DR
Achieving natural-sounding synthetic speech requires precise "phone duration modeling"—knowing exactly how long a sound should last to convey emotions like joy or anger. This study proves that by stripping away irrelevant linguistic noise using Relief and CFS feature selection, we can significantly boost the accuracy of Greek emotional speech models across ten different machine learning architectures, bringing us closer to the "human-like" 20ms error threshold.
The Challenge: The "Noise" in Emotional Prosody
In the world of Text-to-Speech (TTS), prosody is the soul of the voice. It conveys intent and emotion through rhythm and duration. However, modeling this is notoriously difficult because speech features are highly dimensional. When we try to predict how long a "vowel" lasts in a state of Anger vs. Sadness, we often overwhelm models with 90+ variables ranging from syllable position to accentual patterns.
The problem? Many of these features are redundant. This redundancy doesn't just slow down the system; it actively confuses the model, leading to higher Root Mean Square Error (RMSE) and robotic, "flat" emotional delivery.
Methodology: Pruning the Feature Forest
The researchers focused on a Modern Greek emotional database featuring a professional actress expressing five states: Anger, Fear, Joy, Neutral, and Sadness.
1. The Selection Toolkit
They compared two sophisticated algorithms to sift through 93 features:
- Relief: A statistical approach that assigns weights to features based on how well they distinguish between similar instances, aiming for a "maximal margin."
- CFS (Correlation-based Feature Selection): This uses a Genetic Algorithm to find the "sweet spot"—a subset of features that are highly correlated with the target (duration) but have low correlation with each other (minimizing redundancy).
2. The Model Battery
They didn't just test one model; they tested ten, including M5p Model Trees, Linear Regression, and Meta-learning (Bagging) algorithms. This provided a comprehensive look at whether feature selection helps everyone or just specific types of learners.

Insights from the Results
The experimental data (as seen in the tables above) revealed several critical technical insights:
- Consistency is Key: In almost every emotional category (except Joy), using a feature selection algorithm resulted in the lowest error rates.
- The M5p Advantage: The M5p tree was the "MVP." By using a greedy algorithm to build model trees with linear regression functions at the leaves, it managed to handle the non-linear nature of emotional speech better than standard Linear Regression.
- Emotional Sensitivity: Different emotions preferred different selection methods. Relief worked best for "Fear" and "Neutral," while CFS was more effective for "Anger" and "Sadness." This suggests that the phonetic "markers" of fear might be more margin-distinct, while anger features are more structurally correlated.

Deep Insight: Why Why This Matters for TTS
The industry standard for "high quality" is an error under 20ms. Without feature selection, many models hovered around 22-30ms. By introducing Relief and CFS, the researchers pushed models like the AR.M5pR and M5p into the sub-20ms range (e.g., 19ms for Joy and Fear).
This shift isn't just a statistical win; it's a perceptual one. Reducing duration error directly correlates with a reduction in the "uncanny valley" effect of synthetic emotional speech.
Conclusion & Future Outlook
This work confirms that Feature Selection is a vital preprocessing step in the prosody pipeline. While modern TTS has shifted toward neural architectures (Tacotron, FastSpeech), the underlying principle remains: the quality of prosodic output is fundamentally limited by the relevance of the input features.
Future research could investigate whether these selected feature subsets remain consistent across different languages or if "emotional feature importance" is culturally and linguistically specific to Greek.
