Beyond Traditional LBP: Systematic Benchmarking of Facial Features for Emotion Recognition

Feature Extraction and Feature Selection for Emotion Recognition using Facial Expression

2020-09-01
Devashi Choudhary, Jainendra Shukla
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a systematic benchmarking and selection of facial features for Facial Expression Recognition (FER). It proposes a comprehensive framework evaluating 18 distinct handcrafted features (46,352 total dimensions) and discovers that the Bag of Visual Words (BoVW) approach using only 20% of selected features achieves superior performance on the CK+ dataset.

TL;DR

Is more data always better? In the realm of Facial Expression Recognition (FER), this paper argues no. By evaluating 18 types of handcrafted features, the researchers at IIIT-Delhi found that 80% of facial features are noise. By utilizing a "Bag of Visual Words" (BoVW) approach and focusing on a high-performing subset of features like HOG and the newly introduced Grey Co-matrix, they achieved 85.9% accuracy with significantly reduced computational overhead.

The "Small Feature Set" Bottleneck

Despite decades of research into FER, most studies suffer from a "narrow vision" problem—they focus on a single feature type (like Local Binary Patterns) and evaluate it on specific contexts. There has been a lack of a systematic "stress test" to determine which features actually drive performance across the board. Furthermore, high-dimensional data poses a massive challenge for real-time systems like driver monitoring or human-robot interaction.

Methodology: The Battle of Features

The authors didn't just pick a few features; they implemented 18 distinct descriptors, creating a massive feature vector of 46,352 dimensions.

1. Two Structural Approaches

The study compared two distinct pipelines:

  • The Formal Approach: Treats the entire face as a single region for feature extraction.
  • The BoVW Model: Divides the face into local patches (Eyes vs. Nose/Mouth), creates a "visual vocabulary" through K-means clustering, and represents images as histograms of these visual words.

2. Information-Theoretic Feature Selection

To prune the massive feature set, they employed three powerful algorithms:

  • JMI (Joint Mutual Information): Focuses on eliminating redundancy.
  • CMIM (Conditional Mutual Information Maximization): Prefers uncorrelated features.
  • MRMR (Minimum-Redundancy Maximum-Relevance): Balances feature-class correlation with feature-feature distance.

Image_Placeholder Figure 1: The Formal pipeline vs. the superior BoVW pipeline.

Key Discoveries: New Kings of FER

The research yielded several surprising insights that challenge the status quo:

  • The 20% Rule: Increasing features beyond 20% of the total identification set (roughly 9,270 features) does not improve accuracy. In many cases, adding more features actually degrades performance due to the "curse of dimensionality."
  • HOG is Still King: Histogram of Oriented Gradients (HOG) emerged as the most significant feature due to its ability to capture gradient and magnitude orientations of eyes and mouths effectively.
  • The Rise of Grey Co-matrix: Explored for the first time in FER, the Grey Co-matrix (GLCM) outperformed classic benchmarks like LBP. It captures spatial relationships between pixel intensities that traditional filters miss.
  • LDSP over LBP: While Local Binary Patterns (LBP) are famous, Local Directional Position Patterns (LDSP) proved to be much more robust for emotional cues.

Image_Placeholder Figure 2: Relative frequency of feature significance, showing HOG and LDSP as top performers.

Experimental Validation

Using the CK+ (Extended Cohn-Kanade) dataset, the authors tested their assumptions against 8 basic emotions.

  • Formal Approach Accuracy: 84.53%
  • BoVW Model Accuracy: 85.9% (The winner)

The BoVW approach wins because it captures local variations (like a slight twitch in the eye or a curve of the lip) more effectively than a global whole-face analysis.

Critical Insight & Future Outlook

This paper serves as an essential "manual" for anyone designing non-neural network-based FER systems. It proves that handcrafted features are not dead; they are simply under-optimized.

Limitations: The study is confined to the CK+ dataset and discrete emotion labels (e.g., "Happy"). Real-world emotions are often a spectrum (Valence/Arousal), and future work needs to bridge this gap by using multi-dataset training to improve generalization across different lighting and demographic conditions.

Conclusion: If you are building a lightweight FER system, stop using every feature available. Focus on 20%, prioritize HOG and Grey Co-matrix, and utilize a patch-based BoVW architecture for the best results.

Find Similar Papers

Try Our Examples

  • Search for recent papers that compare deep learning-based automated feature extraction with the handcrafted BoVW approach for facial expression recognition.
  • Which original research first established the use of Grey Level Co-occurrence Matrix (GLCM) for texture analysis, and how has its implementation evolved for dynamic facial analysis?
  • Investigate studies that apply the 20% feature selection threshold found in this paper to multi-modal emotion recognition involving both facial and physiological signals.
Contents
Beyond Traditional LBP: Systematic Benchmarking of Facial Features for Emotion Recognition
1. TL;DR
2. The "Small Feature Set" Bottleneck
3. Methodology: The Battle of Features
3.1. 1. Two Structural Approaches
3.2. 2. Information-Theoretic Feature Selection
4. Key Discoveries: New Kings of FER
5. Experimental Validation
6. Critical Insight & Future Outlook