Social Collaborative Filtering: Leveraging Facebook "Likes" to Solve the Cold-Start Crisis
Social collaborative filtering for cold-start recommendations
This paper introduces a generalized matrix algebra framework for the user cold-start recommendation task in online retail. It leverages social side information, specifically Facebook Page Likes, to replace missing interaction data, achieving a 3-fold improvement in Mean Average Precision (mAP) over traditional baselines.
TL;DR
The user cold-start problem is the "day zero" challenge of recommender systems—how do you suggest items to someone you know nothing about? This paper demonstrates that Facebook Page Likes are a goldmine for retail prediction. By redesigning Collaborative Filtering (CF) into a generalized matrix framework, the authors achieved a 300% improvement in recommendation accuracy by treating "likes" as a proxy for latent purchase intent.
Problem & Motivation: The Poverty of Demographics
When a new user arrives at an online store like Kobo, the system is blind. Traditional CF relies on an "Interaction Matrix" (User x Item). If that matrix is empty, the system defaults to "Most Popular" items—a one-size-fits-all approach that ignores individual taste.
While some systems try to use Demographics (Age, Gender, Location), this data is often too coarse. Knowing a user is a "30-year-old male in Australia" doesn't tell you if they prefer Sci-Fi novels or cookbooks. The authors' insight was that social content (what people like on Facebook) provides a much higher resolution of their personality and interests than their demographic profile or their list of friends.
Methodology: Bridging Social and Retail via Matrix Algebra
The core contribution is a generalized matrix algebra framework that decouples the similarity calculation from the recommendation calculation.
The Framework Transition
In standard CF, the system calculates item-item similarity based on co-purchases. In this paper, the authors replace the item dimension with a Personal Attribute (P) dimension.
- Personal Attributes (P): These can be demographics, Facebook friends, or Page Likes.
- Training Phase: The system learns the relationship between attributes and items () from existing users who have both Facebook data and purchase history.
- Recommendation Phase: For a cold-start user, the system only needs their attributes (). The recommendation is generated through .

The authors experimented with various operators for matrix multiplication (), finding that Cosine Similarity for both calculating attribute-item similarity and final scoring (Cos-Cos) yielded the most robust results.
Experiments: Why "Likes" Matter More Than "Friends"
The study used a significant dataset from ebook retailer Kobo:
- 30,000 users.
- 6 million Facebook Page Likes.
- 9 million Facebook Friend connections.
Key Results
The comparison was stark. Page Likes didn't just beat the baseline; they dominated every other side information category.
| Metric | Most Popular | Demographics | Page Likes (This Work) |
|---|---|---|---|
| Precision@1 | 0.006 | 0.014 | 0.038 |
| Recall@10 | 0.074 | 0.083 | 0.173 |
| mAP | 0.026 | 0.035 | 0.075 |
Interestingly, the "Friend Network" performed poorly—almost as badly as basic demographics. This suggests that who you know is far less predictive of your purchasing habits than what you explicitly like.
Figure: The system shows steady gains as more training data is added, though it hits diminishing returns after 80% data saturation.
Critical Analysis & Conclusion
Takeaway
This work validates the "Digital Breadcrumbs" theory: our online interactions (likes) are high-fidelity signals of our latent preferences. By abstracting Collaborative Filtering into a matrix algebra problem, we can effectively "import" knowledge from social networks into retail environments.
Limitations
- Privacy & Access: The method assumes users grant permission to their Facebook data, which is increasingly difficult in the era of GDPR and tightened API access (e.g., Cambridge Analytica aftermath).
- Sparsity: While Page Likes are better than purchase history, they are still sparse. Users with very few likes (under 5) still see poor performance.
- Cross-Domain Noise: Not all "likes" translate to "buys." Some users like pages for irony or news updates, which can introduce noise into a retail book recommender.
Future Outlook
The authors suggest that future iterations could use Collective Matrix Factorization, allowing the system to learn the latent factors of items and social tags simultaneously rather than using simple neighborhood-based similarities. This would likely handle the noise and sparsity even more effectively.
