SOLVER: Bridging the Affective Gap via Scene-Object Interrelated Reasoning
10931_SOLVER: Scene-Object Interrelated Visual Emotion R
The paper proposes SOLVER, a novel Scene-Object interreLated Visual Emotion Reasoning network for Visual Emotion Analysis (VEA). It leverages Graph Convolutional Networks (GCN) to model object-object interactions and a scene-based attention mechanism for scene-object interactions, achieving state-of-the-art performance across eight public datasets.
Executive Summary
TL;DR: Reasoning how an image "feels" is significantly harder than identifying what is "in" an image. The SOLVER (Scene-Object interreLated Visual Emotion Reasoning) network addresses this by shifting the focus from individual objects to the interactions between objects and their environment. By utilizing Graph Convolutional Networks (GCN) and a scene-guided attention mechanism, SOLVER sets a new benchmark in Visual Emotion Analysis (VEA) across eight major datasets.
Background: This work moves VEA from simple classification to a complex reasoning task, positioning itself as a structural bridge between computer vision and cognitive psychology.
The "Affective Gap" and the Motivation
Why can a red rose evoke "Contentment" in a wedding scene but "Sadness" at a funeral? This is the core challenge of VEA. Current SOTA methods often use a "direct mapping" approach—feeding a CNN an image and expecting an emotion label. This ignores the interactionist perspective in psychology: emotions are evoked by the perception of multiple meaningful objects in rich surroundings.
The authors identify two critical missing links in prior work:
- Object-Object Interactions: Multiple objects jointly contribute to the final emotion.
- Scene-Object Interactions: The scene acts as the "emotional tone" that guides how we perceive specific objects.
Methodology: The Architecture of Feeling
SOLVER's architecture is divided into two distinctive branches that mirror human cognitive processes:
1. The Emotion Graph (Object-Object Interaction)
Instead of treating objects as isolated entities, SOLVER builds a graph:
- Nodes: Semantic concepts (via GloVe embeddings) to maintain stability across instances.
- Edges: Emotional affinities calculated from visual features extracted by a Faster R-CNN.
- Reasoning: A 4-layer GCN propagates information across this graph. This allows the model to "reason" that a "bride" and "flowers" together represent a positive event.

2. Scene-Object Fusion (Environmental Guidance)
The model extracts a global scene feature using ResNet-50. Rather than simply concatenating it, the authors propose a Scene-based Attention Mechanism. The scene feature acts as a "query" to determine which objects are most emotionally relevant in that specific context, weighting the fused object features accordingly.

Experimental Mastery
SOLVER was tested against a massive battery of baselines including Fine-tuned ResNet, WSCNet, and MldrNet.
- Superior Performance: On the large-scale FI dataset, it reached 72.33% accuracy, nearly a 5% lead over generic ResNet-50.
- The "Nodal" Sweet Spot: Hyper-parameter analysis revealed that using the top 10 objects (N=10) provides the best balance between informative interaction and noise reduction. Accuracy drops if too many redundant regions are added.

Deep Insights: Interpretability
The authors used TF-IDF and attention weights to visualize "Emotional Object Concepts." For the category "Awe," concepts like mountains, sunset, and horizon received the highest weights. For "Anger," the model learned to focus on human/animal facial features like opened mouths and eyebrows. This confirms that the model is truly learning the stimuli suggested by psychological theory rather than just memorizing background textures.
Critical Analysis & Future Outlook
Takeaway: SOLVER proves that graph-based relational reasoning is essential for high-level cognitive tasks like emotion analysis.
Limitations:
- Human Centricity: In Emotion State Recognition (ESR) tasks—where the goal is to see how a person in the photo feels—SOLVER struggles because it doesn't specifically model facial expressions or body language over general object features.
- Structural Requirements: It struggles with abstract drawings or images containing only a single object/scene where "interaction" is impossible.
Future Work: Combining this scene-object interaction with dedicated facial expression modules could lead to a unified "Theory of Mind" model for AI.
