Instagram as a Mirror: Predicting Personality via Visual vs. Content Features

Predicting Users' Personality from Instagram Pictures: Using Visual and/or Content Features?

2018-07-03
Bruce Ferwerda, Marko Tkalcic, M. Tkalcic
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores personality prediction from Instagram photos using two distinct feature sets: visual features (hue, saturation, brightness) and content features (semantic tags and facial detection). Utilizing the Big Five Inventory (BFI) and a dataset of 54,962 pictures from 193 users, the study demonstrates that both feature types independently serve as effective predictors, with the RBF Network achieving significant improvements over the baseline.

TL;DR

Can your Instagram feed reveal your personality? This study proves it can. By analyzing ~55,000 images, researchers found that visual aesthetics (color/brightness) and image content (what is actually in the photo) are both powerful predictors of the Big Five personality traits. However, contrary to expectations, "more data" isn't always better: combining these two feature sets offers no significant performance boost over using them individually.

Background Positioning

In the landscape of User Modeling, this paper bridges the gap between Computer Vision and Psychology. While prior SOTA work focused on Twitter text or Facebook "Likes," this research targets the highly curated visual world of Instagram, treating the platform as a digital laboratory for behavioral traces.

The Core Conflict: Style vs. Substance

The researchers set out to solve a classic modularity question:

  1. Visual Features (Style): Do extraverts prefer high-saturation, vibrant images? Do neurotic individuals apply specific filters that lower brightness?
  2. Content Features (Substance): Does an "Open" personality post more pictures of architecture and art? Does "Agreeableness" correlate with the number of faces in a photo?

The goal was to determine if these two "channels" contain unique information or if they are simply two sides of the same coin.

Methodology: Decoding the Pixel

The authors employed a dual-pipeline approach to process 54,962 Instagram pictures from 193 participants who underwent BFI (Big Five Inventory) testing.

1. The Visual Pipe

Using the HSV (Hue-Saturation-Value) color space, they mapped physical pixels to psychological states using the PAD Model (Pleasure-Arousal-Dominance).

  • Pleasure:
  • Arousal:
  • Dominance:

2. The Content Pipe

Leveraging the Google Vision API, the team extracted 4,090 unique labels, which were then clustered into 17 semantic categories such as Architecture, Food, Leisure, and Weapons.

Methodology Overview (Note: This study utilizes the RBF Network as the primary high-performing architecture for mapping these features to BFI scores.)

Experimental Insights: RBF Network Dominance

The study compared three classifiers: M5' Rules, Random Forest, and RBF (Radial Basis Function) Networks.

The results (Table 1) show that the RBF Network was the clear winner, achieving the lowest Root Mean Square Error (RMSE) across all personality traits.

TraitBaseline (ZeroR)RBF (Visual)RBF (Content)
Openness0.76190.72310.7133
Conscientiousness0.72010.61750.6375
Agreeableness0.64830.59710.6207

Performance Comparison Table

The Big Surprise: Redundancy in Data Fusion

The most intriguing finding is the failure of feature fusion. Intuitively, combining "how it looks" with "what it is" should provide a 360-degree view of the user. However, the study found that the RMSE for combined features adjusted toward the average.

Why? This suggests a high correlation between visual choice and content choice. For instance, a user seeking to express "Extraversion" might simultaneously choose to post a picture of a crowded party (Content) using high-saturation, high-arousal filters (Visual). The information is redundant.

Critical Analysis & Conclusion

Takeaway

For developers building personality-aware Recommender Systems, this is a "less is more" verdict. You don't need expensive semantic tagging if you have low-level pixel data (or vice-versa).

Limitations

  • Dataset Size: With 134 valid responses, the sample size is modest, which explains why the RBF Network (known for handling small datasets well) outperformed Random Forest.
  • Evolution of Content: The study was conducted in 2018. In 2026, Instagram's shift towards "Reels" (video) and AI-generated content likely requires new feature extraction methodologies.

Future Outlook

The next frontier is moving beyond static images to temporal patterns—how a user's visual style changes over time—as a predictor of psychological well-being or mood shifts.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning embeddings (like CLIP or ResNet) for personality prediction on Instagram to see if they outperform the manual feature extraction used in this 2018 study.
  • Which original research established the Pleasure-Arousal-Dominance (PAD) model for image emotion, and how has its application in personality computing evolved since Valdez and Mehrabian (1994)?
  • Are there any studies investigating cross-platform personality prediction that combine Instagram's visual cues with LinkedIn's professional text to create a more holistic user model?
Contents
Instagram as a Mirror: Predicting Personality via Visual vs. Content Features
1. TL;DR
2. Background Positioning
3. The Core Conflict: Style vs. Substance
4. Methodology: Decoding the Pixel
4.1. 1. The Visual Pipe
4.2. 2. The Content Pipe
5. Experimental Insights: RBF Network Dominance
6. The Big Surprise: Redundancy in Data Fusion
7. Critical Analysis & Conclusion
7.1. Takeaway
7.2. Limitations
7.3. Future Outlook