KAVAN: Decoding Emotions in GIFs through Human-Centered Attention

Human-Centered Emotion Recognition in Animated GIFs

2019-07-01
Zhengyuan Yang, Yixuan Zhang, Jiebo Luo
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the Keypoint Attended Visual Attention Network (KAVAN) for human-centered emotion recognition in animated GIFs. By leveraging a facial attention module and a Hierarchical Segment LSTM (HS-LSTM), the method achieves SOTA performance on the MIT GIFGIF dataset through multi-task learning for both classification and intensity regression.

TL;DR

Animated GIFs have become the lingua franca of digital emotion, yet AI usually treats them as mere "short videos." Researchers from the University of Rochester have introduced KAVAN (Keypoint Attended Visual Attention Network), a framework that prioritizes human facial expressions and hierarchical temporal structures. By using noisy facial keypoints as a guide rather than a crutch, KAVAN achieves new SOTA results on the MIT GIFGIF dataset, proving that in the world of GIFs, the face is the window to the soul.

Problem & Motivation: Why GIFs are Not Just "Small Videos"

GIFs possess two unique properties that traditional Computer Vision models often ignore:

  1. Human-Centricity: Over 50% of GIFs feature human faces, and most others feature personified characters. Traditional CNNs often get distracted by background clutter.
  2. Temporal Density: Unlike long videos, GIFs have no "filler" frames. Every frame is highly salient. Standard LSTMs tend to lose information from early frames by the time they reach the end of the sequence.

Previous attempts to use facial keypoints were brittle—if the keypoint detector failed (common in low-res or stylized GIFs), the whole model failed. The authors set out to create a system that is robust to "missing" data and captures the essence of temporal evolution.

Methodology: The KAVAN Architecture

KAVAN's innovation lies in two distinct modules: the Facial Soft Attention Module and the HS-LSTM.

1. Robust Facial Attention

Instead of feeding keypoint coordinates directly into the network, KAVAN treats them as supervision. The model learns to predict a facial region mask () by comparing its internal attention to a heatmap generated from estimated keypoints.

  • The "Why": If a keypoint detector has low confidence (common in blurry GIFs), the supervision's weight is lowered. This makes the model "smart" enough to find faces even when the keypoint labels are wrong or missing, as seen in cartoon characters.

Model Architecture Fig 1: Overall structure of KAVAN featuring the attention module (blue) and the temporal module.

2. HS-LSTM: Temporal Heirarchy

To solve the "forgetting" problem, the Hierarchical Segment LSTM (HS-LSTM) splits a GIF into segments.

  • Tier 1: Learns coarse representations for each segment.
  • Tier 2: Uses its own frames plus the knowledge from Tier 1 to produce a refined global representation. This ensures that even the very first frame of a 2-second GIF contributes significantly to the final classification.

HS-LSTM Structure Fig 2: A two-tier HS-LSTM learning features from coarse-to-fine resolution.

Experiments & Results: Performance That Matters

The researchers tested KAVAN on the MIT GIFGIF dataset, targeting 17 distinct emotions (e.g., contempt, relief, pride).

Quantitative SOTA

The results show a clear additive benefit of each module:

  • Baseline (ResNet-50 + LSTM): 61.47% Accuracy
  • With Soft Attention: +2.08%
  • With HS-LSTM: +1.08%
  • Multi-Task Learning (MTL): By training for both category classification and emotion intensity regression simultaneously, the model reached 68.27% accuracy.

Qualitative Interpretability

Perhaps most impressively, KAVAN demonstrates "cross-domain" intelligence. In Fig 3 below, we see that even though the keypoint detector might provide poor data for a cartoon character, the Attention Mask (lower row) successfully isolates the character's face, proving the model has learned the semantic concept of a face.

Interpretability Results Fig 3: Visualization of the attention masks. Note how the model focuses on the facial region even in artistic/cartoon styles.

Critical Insight & Conclusion

KAVAN proves that soft supervision is often superior to hard feature fusion. By teaching the model where to look (using keypoints as a hint) rather than what the keypoints are, the authors created a robust system that handles the chaotic variety of social media content.

Future Outlook: While KAVAN is localized to GIFs, its hierarchical temporal approach could be revolutionary for other "dense" video tasks, such as sign language recognition or surgical video analysis, where every millisecond counts.

Find Similar Papers

Try Our Examples

  • Find recent papers that apply facial keypoint-guided attention to video-based affective computing or sentiment analysis.
  • Which study first introduced the Circumplex Affect Model, and how has it been adapted for multi-task learning in deep emotion recognition?
  • Explore if Hierarchical Segment LSTM (HS-LSTM) architectures have been successfully applied to other short-form video tasks like action recognition or micro-expression detection.
Contents
KAVAN: Decoding Emotions in GIFs through Human-Centered Attention
1. TL;DR
2. Problem & Motivation: Why GIFs are Not Just "Small Videos"
3. Methodology: The KAVAN Architecture
3.1. 1. Robust Facial Attention
3.2. 2. HS-LSTM: Temporal Heirarchy
4. Experiments & Results: Performance That Matters
4.1. Quantitative SOTA
4.2. Qualitative Interpretability
5. Critical Insight & Conclusion