Beyond the Surface: Leveraging 3D Action Units for Enhanced Emotion Recognition
Facial Expression Based Emotion Recognition Using Neural Networks
2018-01-01
Summary
Problem
Method
Results
Takeaways
Abstract
This paper presents a facial emotion recognition system that classifies seven emotional states using 17 Action Units (AUs) tracked by the Microsoft Kinect v2 sensor. By employing an Artificial Neural Network (ANN) with scaled conjugate gradient backpropagation, the authors achieve a high subject-dependent classification accuracy for real-time human-computer interaction applications.
## TL;DR
Researchers developed an emotion recognition system using the **Kinect v2 sensor** and **Artificial Neural Networks (ANNs)**. By tracking 17 distinct 3D Facial Action Units (AUs), the system achieves a remarkable **95.8% accuracy** in subject-dependent scenarios, though it reveals clear limitations when generalizes to new individuals or across different genders.
## Executive Summary
In the landscape of Human-Computer Interaction (HCI), understanding human emotion is the "Holy Grail." While traditional 2D computer vision has dominated the field, it often falters under varying lighting conditions. This paper shifts the focus to 3D depth data, utilizing the Microsoft Kinect v2 to track muscle movements in real-time. By expanding the feature set from the standard 6 AUs to 17, the authors aim to capture the subtle nuances of joy, sadness, surprise, anger, fear, disgust, and neutral states.
## The Motivation: Why 3D and Why More Action Units?
Most existing emotion recognition frameworks rely on 2D images, which are computationally heavy and sensitive to the environment. The authors argue that:
1. **Depth Matters**: 3D sensors provide spatial coordinates that are invariant to certain lighting shifts.
2. **Granularity is Key**: Previous studies using Kinect v1 only utilized 6 AUs (mostly upper face). By utilizing 17 AUs, this research captures muscle activities around the mouth and jaw—critical areas for distinguishing emotions like *disgust* and *joy*.
## Methodology: The ANN Architecture
The core of the system is a feed-forward Neural Network.
- **Input Layer**: 17 features representing the displacement and weight of Action Units (AUs).
- **Hidden Layer**: 10 neurons with a Sigmoid activation function.
- **Output Layer**: 7 softmax-style outputs representing the emotional states.

*Fig 1: The Neural Network structure used for classification.*
The training utilized the **Scaled Conjugate Gradient Backpropagation** algorithm, known for its efficiency in handling network weights without requiring extensive manual parameter tuning.
## Experiments and The Generalization Gap
The study involved six subjects (3 male, 3 female) performing emotions in a controlled experimental setup.
### Key Results:
- **Subject Dependent**: 95.8% Accuracy. When the model "knows" the person's face structure, it is nearly flawless.
- **Subject Independent**: 67.03% Accuracy. Testing the model on a person it has never seen before leads to a significant performance drop.
- **Gender Sensitivity**: 56% Accuracy. Training only on males and testing on females resulted in the lowest performance, suggesting that the "Action Unit" signatures for the same emotion differ significantly across genders.

*Table 1: Classification results for unseen test subjects.*
## Critical Analysis & Conclusion
This research successfully demonstrates that **17 Action Units** provide a robust feature set for 3D emotion recognition. The high subject-dependent accuracy suggests these features are highly discriminative.
**However, the "Generalization Gap" is the elephant in the room.** The drop from 95% to 67% indicates that the model is likely over-fitting to the specific facial morphologies of the training subjects. The gender-based test further proves that "one size does not fit all"—a smile or a scowl manifests differently across different demographic profiles.
### Future Outlook
To move from a laboratory success to a consumer product, future research must:
- Incorporate **Principal Component Analysis (PCA)** to isolate the most significant features.
- Utilize **Fusion Algorithms** that combine AUs with Feature Point Positions (FPPs) to provide a more holistic view of the face.
- Expand datasets to include higher diverse age groups and ethnicities to resolve the demographic bias.
