Beyond Regions: Improving Image Emotion Recognition via Weakly Supervised Intensity Learning

Weakly Supervised Emotion Intensity Prediction for Recognition of Emotions in Images

2020-07-08
Haimin Zhang, Min Xu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel end-to-end deep neural network for image emotion recognition that leverages weakly supervised emotion intensity learning. By integrating an Intensity Prediction Stream built on a Feature Pyramid Network (FPN) with two classification streams, the method achieves new SOTA results on benchmarks like FI-8 and Emotion-6.

TL;DR

Recognizing emotions in images is inherently difficult because emotions are subjective and rarely occupy the entire frame. This paper introduces an end-to-end framework that doesn't just classify an image, but predicts an Emotion Intensity Map to pinpoint exactly where the "feeling" comes from. By using a weakly supervised approach that transforms class activation into intensity guidance, the authors achieved SOTA performance on the FI-8 and Emotion-6 datasets.

The "Weak Label" Bottleneck

In standard image classification (like ImageNet), a "cat" is usually a concrete object. In emotion recognition, a "sad" label might only apply to a small drooping flower in a vast neutral landscape. Most existing datasets only provide image-level labels, which are "weak" because they don't tell the model where to look.

Prior attempts to solve this used:

  1. Handcrafted features: Limited by the expert's imagination.
  2. Region proposals: Computationally heavy and often required manual bounding boxes for pre-training.

The authors' insight was simple yet powerful: Emotion is a continuous intensity field, not just a binary box.

Methodology: The Triple-Stream Architecture

The proposed network consists of three distinct yet cooperative streams:

  1. The Probing Stream (First Classification): A standard CNN that generates initial class activation maps (CAM). These maps serve as "pseudo-ground truth" for what regions drive the emotion.
  2. The Discovery Stream (Intensity Prediction): Built on a Feature Pyramid Network (FPN), this stream learns to predict the intensity map directly from the image. It uses a combination of three losses:
    • RMSEL: For pixel-wise intensity accuracy.
    • Gradient Loss: To keep the edges of emotional regions sharp.
    • Surface Normal Loss: To ensure the geometric "shape" of the intensity reflects the image structure.
  3. The Refinement Stream (Second Classification): This stream takes the original features and "multiplies" them by the predicted intensity map. This forces the model to ignore background noise and focus its representation on high-intensity emotional stimuli.

Model Architecture Figure: The three-stream architecture showing the flow from Pseudo-Intensity generation to final classification.

Why the FPN Matters

By using an FPN (shown below), the network can extract multilevel features. This is critical because some emotional cues are small (a subtle facial expression), while others are global (the color palette of a sunset). The FPN's top-down pathway combines these semantic scales into a single, high-resolution intensity map.

FPN Detail Figure: The Intensity Prediction Subnetwork built on Top of FPN.

Experimental Results & Insights

The results across datasets were consistently superior to vanilla architectures and previous region-based SOTA like Rao et al.

  • FI-8 Dataset: Jumped from 66.16% (Vanilla ResNet-101) to 75.91%.
  • Emotion-6: Achieved 60.41%, noticeably better than much larger models like ResNet-152.

One fascinating takeaway from the ablation studies is that the "Surface Normal Loss" and "Gradient Loss" (typically used in depth estimation) significantly helped. This suggests that the spatial "contours" of an emotion are just as important as the raw pixel values.

Visual Results Figure: Visualization comparing CAM-generated maps (left) vs. the model's Predicted Intensity Maps (right).

Critical Analysis & Future Outlook

While the method is robust, the confusion matrices reveal that "Fear" and "Anger" remain difficult to distinguish—often because these emotions share similar visual triggers in the wild.

Future Directions:

  • Cross-Modal Guidance: Could textual metadata (tags/comments) be used to further refine the pseudo-intensity maps?
  • Temporal Stability: Applying this intensity-based logic to video clips, where emotion intensity fluctuates over time.

In conclusion, this work proves that we don't need expensive manual annotations to understand "where" an emotion is. By teaching a network to predict its own attention maps, we get a model that is both more accurate and more interpretable.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Feature Pyramid Networks (FPN) or attention mechanisms specifically for weakly supervised affective computing in images.
  • Which study first introduced the Class Activation Mapping (CAM) technique for localization, and how have subsequent works adapted it for non-object detection tasks like emotion recognition?
  • Explore research that applies emotion intensity prediction or similar spatial weighting methods to video-based emotion recognition or multi-modal sentiment analysis.
Contents
Beyond Regions: Improving Image Emotion Recognition via Weakly Supervised Intensity Learning
1. TL;DR
2. The "Weak Label" Bottleneck
3. Methodology: The Triple-Stream Architecture
4. Why the FPN Matters
5. Experimental Results & Insights
6. Critical Analysis & Future Outlook