Beyond Pixels: Landmark-Based Robustness in Emotion Recognition

An Adversarial Attacks Resistance-based Approach to Emotion Recognition from Images using Facial Landmarks

2021-04-14
Harisu Abdullahi Shehu, William Browne, Hedwig Eisenbarth
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a facial landmark-based approach for Emotion Recognition (ER) designed specifically to resist adversarial attacks. By utilizing Dlib for landmark extraction and a Random Forest (RF) classifier, the method achieves robustness comparable to ResNet on clear data while significantly outperforming it under adversarial perturbations.

TL;DR

Deep learning models are notoriously fragile—a tiny bit of noise can flip a "Happy" prediction to "Angry." This paper demonstrates that by moving away from raw pixel-based classification and toward Facial Landmark Geometry, we can build emotion recognition systems that are virtually immune to common adversarial attacks (FGSM) and occlusions, while being 100x faster than traditional CNNs like ResNet.

Problem & Motivation: The "Homogeneous" Trap

Current SOTA models for emotion recognition often treat the task as a standard image classification problem. They learn the color distribution and pixel gradients (homogeneous knowledge) rather than the underlying anatomy of an expression.

The authors argue that this makes models susceptible to:

  1. Adversarial Attacks: Infinitesimal noise that shifts pixel values but doesn't change the expression to a human.
  2. Environmental Variance: Changes in skin color, lighting, or accessories (like glasses) that shouldn't affect emotional state analysis.

Inspired by Ekman’s Facial Action Coding Systems (FACS), the researchers hypothesize that the relative movement of facial landmarks is the true signal, and everything else is just noise.

Methodology: The Power of Relative Geometry

The core of the proposed method is to strip away the "distraction" of pixels.

1. Feature Extraction

Using the Dlib library, the team extracts (x, y) coordinates for key facial features (eyes, eyebrows, nose, mouth, jaw).

2. Spatial Normalization (The Central Point)

Directly feeding raw coordinates into a classifier fails because people appear at different positions in a frame. To solve this, the authors compute a Central Point (C)—the average of all detected coordinates. They then calculate the Euclidean distance of every landmark relative to C.

Sample distance of landmark coordinates Fig. 1: Normalizing landmarks by calculating distances relative to a central point.

3. Classification

Instead of a heavy ResNet, they use a Random Forest (RF). Since the feature space is now reduced to a streamlined vector of distances, the computational overhead vanishes.

Adversarial Testing: ResNet vs. Geometry

The authors subjected both a ResNet and their Landmark-RF model to three types of attacks:

  • Type A/B: Physical occlusions (black boxes over eyes or jaw).
  • FGSM (Fast Gradient Sign Method): Mathematical noise designed to maximize model loss.

FGSM Attack Examples Fig. 2: FGSM attack on ResNet. The human sees no difference, but the model's confidence collapses or predicts the wrong emotion.

Experimental Results

The findings were stark. On the CK+ dataset:

  • Clean Data: Both ResNet and Landmarks achieved ~97% accuracy.
  • Under Attack: ResNet’s accuracy plummeted to 80.8% (Type A) and 90.8% (Type B).
  • The Proposed Method: Maintained an astonishing 96.8% - 97% accuracy, showing almost zero sensitivity to the noise.
Attack TypeResNet AccuracyProposed (Landmark) Accuracy
None97.43%97.14%
Type A (Occlusion)80.86%96.86%
FGSM (Noise)90.00%98.00%

Efficiency Gains

Perhaps the most "product-ready" takeaway: while ResNet took over 5 hours to process the data on a GTX 1080ti, the landmark approach finished in under 2 minutes.

Critical Analysis: Why This Works

The "Secret Sauce" lies in the information bottleneck. By forcing the model to only "see" landmark coordinates, the authors have effectively:

  1. Filtered the Noise: FGSM targets pixel gradients. If the model doesn't look at pixels, the attack has no vector.
  2. Invariance to Color: Histograms (as seen in the paper's Fig. 7) show that while occlusions destroy color distribution, the geometric pattern of a surprise or disgust expression remains intact.

Conclusion & Future Work

This paper is a strong reminder that "Deep Learning" isn't always the only answer. By incorporating domain-specific structural knowledge (Facial Landmarks), we can achieve Adversarial Robustness that pure pixel-crunching models currently lack.

The next hurdle? Testing this on "in-the-wild" datasets where landmark detection itself might be compromised by extreme head poses. If the landmark extractor falls, the system falls—making the robustness of the extractor itself the new frontier for research.

Find Similar Papers

Try Our Examples

  • Search for recent papers that combine facial landmark geometry with Graph Convolutional Networks (GCNs) to improve adversarial robustness in emotion recognition.
  • Which paper first established the Facial Action Coding System (FACS) as a standard for landmark-based emotion analysis, and how does the Dlib implementation utilized here differ from original FACS manual encoding?
  • Investigate state-of-the-art methods that apply adversarial training specifically to landmark detection modules rather than end-to-end pixel classifiers.
Contents
Beyond Pixels: Landmark-Based Robustness in Emotion Recognition
1. TL;DR
2. Problem & Motivation: The "Homogeneous" Trap
3. Methodology: The Power of Relative Geometry
3.1. 1. Feature Extraction
3.2. 2. Spatial Normalization (The Central Point)
3.3. 3. Classification
4. Adversarial Testing: ResNet vs. Geometry
5. Experimental Results
5.1. Efficiency Gains
6. Critical Analysis: Why This Works
7. Conclusion & Future Work