C-GCN: Decoding Human Emotion via Inter-Video Correlations

11137_C-GCN Correlation Based Graph Convolutional Network for Audio-Video Emotion Recognition.

Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces C-GCN, a Correlation-based Graph Convolutional Network for audio-video emotion recognition. It leverages multi-head attention to model inter-class and intra-class relationships between video clips, achieving SOTA results on AFEW (63.55%) and eNTERFACE 05 (97.07%).

TL;DR

Recognizing emotions from audio-visual data is notoriously difficult due to "weak expressions" and "conflicting emotional states." While most researchers focus on fusing audio and video within a single clip, C-GCN (Correlation-based Graph Convolutional Network) shifts the paradigm. By treating video clips as nodes in a global graph and using Multi-Head Attention to learn the edges, the model "consults" other samples to refine its own predictions, achieving state-of-the-art accuracy with remarkable inference speed.

The Missing Piece: Context Beyond the Clip

In traditional Automatic Emotion Recognition (AER), the model is an island. It looks at Video A and tries to decide if the person is "Sad" or "Neutral" based solely on the pixels and audio waves of Video A.

However, emotions are abstract and relative. The authors of C-GCN argue that inter-video correlations—how Video A relates to Video B (even if they are from different people)—are crucial. Current methods suffer because:

  1. They ignore global information across the dataset.
  2. Subtle "weak emotions" (like disgust or surprise) are often misclassified because the model lacks a comparative reference within the feature space.

Methodology: Building the Emotion Graph

The C-GCN pipeline follows a sophisticated tripartite architecture:

1. Robust Feature Extraction

The model uses a dual-flow system:

  • Audio Flow: Spectrograms are processed via a Fully Convolutional Network (FCN) with an attention mechanism to focus on informative time-frequency units.
  • Image Flow: Faces are tracked using dlib, features extracted via FR-Net-B (fine-tuned on FER2013), and temporal dynamics captured via BiLSTM.
  • Fusion: The audio and visual vectors are integrated using Factorized Bilinear Pooling (FBP), which captures complex associations more effectively than simple concatenation.

2. Graph Generation with Multi-Head Attention

This is the core innovation. Instead of a static graph, C-GCN uses:

  • Initial Edges: Based on category labels (for training) and cosine similarity (for testing).
  • Multi-Head Attention: This transforms the sparse initial graph into a set of fully connected edge-weighted graphs. This step eliminates "isolated nodes" and allows the model to predict hidden relationships between different emotion classes.

C-GCN Architecture Fig 1. The overall architecture of C-GCN, illustrating the flow from raw data to graph-based classification.

3. Feature Updating via Dense GCN

To prevent the vanishing gradient problem and the loss of original features, the model employs Densely Connected GCNs. Each layer receives the concatenation of all previous layers' outputs. The information from multiple attention heads is finally fused via mean-pooling to create a highly discriminative video descriptor.

Experimental Results: SOTA Performance

C-GCN was tested on the AFEW and eNTERFACE 05 datasets.

  • Efficiency: Despite the graph complexity, C-GCN reached a classification speed of 0.12s per video, outperforming competitors like Zhou et al. and Hu et al.
  • Accuracy: On AFEW, it achieved 63.55%, an impressive feat for a single-model approach compared to multi-model ensembles.
  • The "Weak Emotion" Breakthrough: The model showed significant improvements in recognizing "Disgust" and "Surprise," which are traditionally the hardest categories to crack.

Performance Comparison Table 1. Ablation study showing the impact of FBP and Multi-Head Attention on final accuracy.

Critical Insight: Why Does It Work?

The genius of C-GCN lies in how it handles Manifold Learning. By using Multi-Head Attention to predict edges, the model essentially learns the underlying structure of the "emotion manifold." If Video A is a "weak" version of Happy, the graph creates a strong edge to Video B (a "strong" Happy), allowing the features of Video B to "pull" Video A into the correct classification zone during the GCN message-passing stage.

Conclusion & Future Outlook

C-GCN successfully demonstrates that "no video is an island." By leveraging the hidden correlations between different samples, the model achieves a more nuanced understanding of affect.

Future Work: The authors suggest the next step is modeling cross-modal correlations—using GCNs to map the interaction between audio features and visual landmarks directly, potentially further reducing the confusion between similar emotional states.


Keywords: Emotion Recognition, GCN, Multi-modal Fusion, Multi-head Attention, Deep Learning.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Convolutional Networks (GCNs) for cross-sample relationship modeling in multi-modal emotion recognition.
  • Which paper first introduced Factorized Bilinear Pooling (FBP) for audio-visual fusion, and how does the current work refine that feature integration method?
  • Explore the application of multi-head attention mechanisms in generating dynamic graph structures for tasks beyond emotion recognition, such as video action recognition or medical image analysis.
Contents
C-GCN: Decoding Human Emotion via Inter-Video Correlations
1. TL;DR
2. The Missing Piece: Context Beyond the Clip
3. Methodology: Building the Emotion Graph
3.1. 1. Robust Feature Extraction
3.2. 2. Graph Generation with Multi-Head Attention
3.3. 3. Feature Updating via Dense GCN
4. Experimental Results: SOTA Performance
5. Critical Insight: Why Does It Work?
6. Conclusion & Future Outlook