SOLVER: Decoding the Relational Language of Visual Emotions

10931_SOLVER: Scene-Object Interrelated Visual Emotion R

Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces SOLVER, a Scene-Object interreLated Visual Emotion Reasoning network for Visual Emotion Analysis (VEA). It utilizes an Emotion Graph with Graph Convolutional Networks (GCN) to model object-object interactions and a scene-based attention mechanism to integrate scene-object relationships, achieving SOTA performance across eight public datasets.

TL;DR

Human emotions are rarely triggered by a single object in isolation. Instead, our feelings are evoked by how objects interact with each other and their environment. SOLVER (Scene-Object interreLated Visual Emotion Reasoning network) is a novel architecture that mimics this cognitive process by using Graph Convolutional Networks (GCN) and Scene-based Attention to bridge the "affective gap" in image analysis.

The Problem: The Affective Gap

In the field of Visual Emotion Analysis (VEA), most deep learning models treat an image as a bag of pixels or a collection of isolated regions. They map these features directly to emotion labels like "Happy" or "Sad." However, psychology tells us that emotion is contextual.

For example, a "red rose" might evoke Contentment at a wedding but Sadness at a funeral. The object is the same, but the interaction with other objects and the scene changes the emotional output. Previous models struggled with this nuance, leading to a performance bottleneck known as the affective gap.

Methodology: Reasoning over Graphs and Scenes

SOLVER addresses this by treating an image as a structured system of interactions.

1. The Emotion Graph (Object-Object Interaction)

The model first detects objects using Faster R-CNN. It then builds an Emotion Graph where:

  • Nodes: Represent semantic concepts (e.g., "dog," "mountain") using GloVe embeddings.
  • Edges: Represent the emotional relationship (affinity) between these objects, calculated using visual features in an emotional embedding space.

By applying Graph Convolutional Networks (GCN), the model allows object features to "communicate" with their neighbors. A "rose" node learns about the "bride" node nearby, enhancing its feature representation with relational context.

SOLVER Architecture

2. Scene-Object Fusion (Scene-Object Interaction)

While objects provide local triggers, the scene sets the global "tone." SOLVER uses a Scene-Object Fusion Module. It doesn't just concatenate features; it uses the global scene feature as a "query" to weight the importance of different objects. This scene-based attention ensures that if the scene is a "stadium," the model pays more attention to "players" and "crowds" to infer excitement.

Scene-Object Fusion Process

Experimental Performance

The researchers tested SOLVER against a battery of SOTA methods (including WSCNet and MldrNet) across 8 datasets.

  • Large-scale datasets: Achieved a significant lead (e.g., 72.33% on FI, 86.20% on Flickr).
  • Ablation Study: The results showed that adding the Emotion Graph + Fusion module improved performance significantly over using just a standard ResNet-50 backbone (which scored 67.53% on FI).

SOTA Comparison Table

Deep Insight: Interpretable AI

One of the most impressive aspects of the SOLVER paper is its interpretability. By visualizing Weighted Frequencies, the authors show which concepts actually drive specific emotions.

  • Awe is consistently linked to "mountains," "cliffs," and "horizons."
  • Excitement correlates with "surfboards," "rafts," and "microphones."

The attention maps demonstrate that the model "looks" at the most emotionally relevant interactions—such as the bared teeth of a leopard when predicting "Anger"—rather than just the animal's body.

Emotional Concept Visualization

Conclusion & Limitations

SOLVER proves that relational reasoning is the key to mastering high-level cognitive tasks like emotion analysis. However, the authors honestly note a limitation: the model currently misses micro-expressions (facial cues) and body language, which are vital for human-centric datasets like LUCFER. Future iterations will likely integrate these "human-centric" features into the existing scene-object graph to create a truly holistic emotional AI.

Find Similar Papers

Try Our Examples

  • Search for recent papers in Visual Emotion Analysis that utilize Graph Neural Networks or Scene Graphs to model contextual relationships.
  • Who first proposed the use of Adjective Noun Pairs (ANPs) for visual sentiment, and how does the Emotion Graph in SOLVER improve upon that semantic approach?
  • Explore if the SOLVER framework for scene-object interaction can be adapted for Emotion Recognition in Context (ERC) tasks involving video or multi-modal data.
Contents
SOLVER: Decoding the Relational Language of Visual Emotions
1. TL;DR
2. The Problem: The Affective Gap
3. Methodology: Reasoning over Graphs and Scenes
3.1. 1. The Emotion Graph (Object-Object Interaction)
3.2. 2. Scene-Object Fusion (Scene-Object Interaction)
4. Experimental Performance
5. Deep Insight: Interpretable AI
6. Conclusion & Limitations