Deciphering the Visual Soul: Personality Inference via Convolutional Neural Networks

Computer Vision and Image Understanding

2022-01-01
Olivier Pradelle, Raphaelle Chaine, David Wendland, Julie Digne
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a deep learning framework for personality inference based on user-curated image collections (Flickr "faves"). By leveraging fine-tuned Convolutional Neural Networks (CNNs) like AlexNet and VGG-16, the authors predict both self-assessed and attributed Big Five personality traits, achieving SOTA performance in social signal processing.

TL;DR

Can the images you "like" on social media reveal who you truly are? This paper demonstrates that deep learning can decode your personality traits—collectively known as the "Big Five"—simply by analyzing your favorite photos. By fine-tuning CNNs, the researchers moved beyond simple color histograms to identify the complex "visual subtexts" that signal traits like Extraversion or Neuroticism.

Background: From Pixels to Personalities

For decades, image understanding focused on what is in a picture (object recognition) or how pretty it is (computational aesthetics). This work defines a third way: Social Profiling. The central hypothesis is that our aesthetic preferences are not random; they are manifestations of our internal psychological makeup. Using the PsychoFlickr corpus, which pairs 60,000 "fave" images from 300 users with their actual personality scores, the authors set out to build an automated "social lens."

The Core Motivation: Why Deep Learning?

Previous SOTA methods used "hand-crafted" features—engineers would manually define rules for brightness or the "rule of thirds." However, personality is nuanced. A person high in Openness might prefer surrealist art, while someone high in Conscientiousness might gravitate toward clean, geometric architecture. CNNs excel here because they learn entangled attributes: they don't just see "blue"; they see the "melancholy of a vast, empty sky."

Methodology: Tuning the Artificial Brain

The researchers utilized high-performance architectures including AlexNet and VGG-16 (VeryDeep-16). The process involved:

  1. Fixed Convolutional Bases: Keeping the early layers (which detect edges and textures) fixed to preserve general visual knowledge.
  2. Fine-Tuning: Modifying the final layers to classify images into "High" or "Low" scores for the Big Five traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism).
  3. Visualization: Using a deconvolution strategy to "see" what the network sees.

Model Architecture and Process

Cracking the Code: What Do the Traits Look Like?

The most fascinating part of the study is the qualitative discovery of trait-specific visual patterns.

  • Neuroticism: High-scorers preferred images with empty regions, faded colors, and "twisted" or "disturbed" shapes. In contrast, low-scorers (more emotionally stable) preferred natural scenery and vibrant water patterns.
  • Extraversion: As expected, high-scorers displayed a clear preference for crowded scenes and people. Low-scorers (introverts) favored indoor settings, macro-photography of plants, and solitary objects.
  • Conscientiousness: This trait was strongly associated with "orderly" visual content—buildings with sharp lines, defined horizons, and symmetrical compositions.

Neuroticism Visual Archetypes Fig: Visual contrast between Low (left) and High (right) Neuroticism preferences.

Experimental Results: Machines vs. Human Perception

The study distinguished between Self-Assessed (how I see myself) and Attributed traits (how a stranger sees me based on my photos).

  • Attributed Traits were easier to predict: The CNN achieved significantly higher accuracy here (~10% improvement). This suggests that we project a "visual persona" that is highly consistent and recognizable to both crowds and AI, even if it slightly differs from our messy internal self-perception.
TraitSelf-Assessed AccuracyAttributed Accuracy
O penness0.530.62
C onscientiousness0.550.67
N euroticism0.520.69

Critical Insight & Future Outlook

The "wisdom of the crowds" is effectively captured within the weights of a CNN. This work proves that images are not just media objects; they are social signals.

Limitations: The dataset is relatively small (300 users), and the binary classification (High vs. Low) simplifies the continuous nature of human personality. Future Work: Moving toward Regression models (predicting exact scores) and integrating Temporal data (how preferences change over time) could turn these tools into powerful assets for clinical psychology and personalized recommendation engines.

Conclusion

By looking at the world through the "eyes" of a fine-tuned VGG-16, we gain a new perspective on human nature. We are, quite literally, what we like.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend personality inference from static images to short-form video preferences (e.g., TikTok or Reels) using multimodal transformers.
  • Which study first introduced the PsychoFlickr dataset, and how have subsequent works improved the labeling consistency between self-assessment and observer attribution?
  • Examine how the visual patterns identified in this paper (like "chaotic" vs. "orderly" compositions) align with established theories in visual psychology such as the Brunswick Lens Model.
Contents
Deciphering the Visual Soul: Personality Inference via Convolutional Neural Networks
1. TL;DR
2. Background: From Pixels to Personalities
3. The Core Motivation: Why Deep Learning?
4. Methodology: Tuning the Artificial Brain
5. Cracking the Code: What Do the Traits Look Like?
6. Experimental Results: Machines vs. Human Perception
7. Critical Insight & Future Outlook
7.1. Conclusion