Beyond Categories: Leveraging Social-Centric Scores for Personalized Venue Suggestion

Venue Suggestion Using Social-Centric Scores

2020-01-01
Mohammad Aliannejadi, Fabio Crestani
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a personalized venue suggestion framework that leverages "Social-Centric Scores" derived from Location-Based Social Networks (LBSNs) like Yelp and Foursquare. By combining review-based sentiment classification and Foursquare "taste keywords" with Learning to Rank (LTR) techniques, the method achieves SOTA performance on the TREC Contextual Suggestion track.

TL;DR

Recommending the perfect restaurant or tourist attraction is difficult because a "thumbs up" doesn't tell the whole story. This paper by Aliannejadi and Crestani moves beyond simple categories (like "Italian Restaurant") by mining social-centric information—specifically user reviews and "taste keywords" from LBSNs. By treating user preference as a classification problem and applying Learning to Rank (LTR), they demonstrate that what other people say about a venue is a far stronger signal than the venue's official classification.

The "Why" vs. the "What": The Motivation

Most recommender systems rely on Collaborative Filtering (CF). While effective, CF hits a wall when data is sparse—which is almost always the case in geography-based tasks. If you haven't visited a city before, the system has no "history" to match you with others.

The authors argue that existing systems focus too much on the "What" (e.g., "The user likes Pizza places") and ignore the "Why" (e.g., "The user likes quiet, family-friendly spots with good wine"). Social networks like Foursquare and Yelp provide the missing link through:

  1. Taste Keywords: Crowdsourced tags like "cozy," "late-night," or "scenic views."
  2. Unstructured Reviews: Textual explanations of the user experience.

Methodology: Engineering the Social Signal

The authors' framework is built on three core pillars:

1. The Frequency-based Score (Keywords)

The system builds a Positive/Negative Keyword Profile for each user. If you frequently rate venues positively that are tagged with "live music," that keyword gains weight in your profile.

2. The Review-based Score (SVM Classification)

Because users don't always leave reviews, the authors use a clever proxy: they train a Linear SVM Classifier using reviews from other users who gave similar ratings to the venues in the target user's history. This models the user's "filter"—how they might react to reading comments about a new place.

3. Learning to Rank (The Fusion)

Instead of just adding these scores together, the authors test several LTR techniques (MART, LambdaMART, RankNet). RankNet, a neural-network-based approach, proved most effective at finding the optimal non-linear combination of these social features.

Model Table Comparison Table 2: Performance comparison showing LTR-S (Social-only) outperforming LTR-All and other SOTA methods.

Experimental Insights: Quality over Quantity

The results from the TREC Contextual Suggestion Track yielded several counter-intuitive findings:

  • Social > Content: The model using only Social scores (LTR-S) outperformed the model using all features (LTR-All). Including venue categories actually introduced "noise" that degraded the ranking.
  • The Power of Pruning: You don't need all reviews. The authors found that using only the most recent reviews or reviews from the most active users actually improved performance. This suggests that the "freshness" and "credibility" of social data are as important as the data itself.

Review Impact Graph Figure 2: Analysis showing how selecting reviews by "Recent" or "Active" users leads to better nDCG@5 scores compared to random selection.

Critical Analysis & Takeaways

Why does it work? The success of this method lies in its ability to capture latent dimensions of preference. A "Pizza Place" (Category) might be a fast-food joint or a high-end date spot. Keywords and reviews distinguish between these, effectively performing a "dimensionality reduction" on the complex space of human taste.

Limitations:

  1. Computation Cost: Training an SVM classifier per user is feasible for 211 users (the study size), but would require significant optimization for a platform like Yelp with millions of users.
  2. Platform Dependency: The method relies on high-quality crowdsourced tags (Foursquare Tastes). Without these, the keyword score performance would likely drop.

Conclusion: This work proves that in the era of LBSNs, the "social crowd" is the best sensor for personalization. Future work involving Word Embeddings (Word2Vec/BERT) could further enhance this by understanding that "cozy" and "intimate" are semantically the same, further smoothing the data sparsity gap.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) to extract "venue taste keywords" or sentiment features for point-of-interest recommendation.
  • Which paper first introduced the RankNet algorithm for Information Retrieval, and how does the pairwise loss function differ from the listwise approach used in ListNet?
  • Explore how social-centric scores or review-based classifiers have been applied to multi-modal recommendation systems beyond the location-based social network domain.
Contents
Beyond Categories: Leveraging Social-Centric Scores for Personalized Venue Suggestion
1. TL;DR
2. The "Why" vs. the "What": The Motivation
3. Methodology: Engineering the Social Signal
3.1. 1. The Frequency-based Score (Keywords)
3.2. 2. The Review-based Score (SVM Classification)
3.3. 3. Learning to Rank (The Fusion)
4. Experimental Insights: Quality over Quantity
5. Critical Analysis & Takeaways