Crowdsourcing the "Subjective" Map: Bridging the Gap in Location-Based Search
Crowdsourcing location-based queries
The paper introduces a framework for crowdsourcing location-based queries using Twitter as a middleware and Foursquare for expert identification. It focuses on non-factual, subjective queries where traditional search engines fail, achieving a 75% answer rate for such queries compared to Google's 29%.
TL;DR
Traditional search engines are great at finding "hotels in Miami," but they fail miserably when you ask for a "cheap, baby-friendly hotel near Broadway." This paper addresses this gap by crowdsourcing location-based queries to Foursquare experts via Twitter. The result? A system that answers 75% of subjective queries where Google manages less than 30%.
Problem & Motivation: The Failure of Factual Search
When we use a smartphone, our queries are often "non-factual." They are relative, multi-dimensional, and highly subjective. The authors discovered that 63% of location-based queries on Twitter fall into this category.
While algorithms can index static web pages, they cannot "feel" the vibe of a coffee shop or know if an apartment allows dogs based on a simple keyword search. The authors' insight is simple: Humans are the best sensors for subjective data. By identifying people who actually visit these locations (using Foursquare check-ins), we can route questions to "local experts" who are more likely to provide high-quality, nuanced answers.
Methodology: Routing Questions to the Right Crowd
The system architecture consists of a sophisticated pipeline designed to filter noise and ensure quality:
- Question Collector: Monitors Twitter for location-specific questions using a template-based approach (e.g., "Anyone... [location]... ?").
- Validator: A human-in-the-loop stage where moderators categorize the question (e.g., Food, Nightlife) and rank its quality.
- Asker (The Expert Finder): This is the "secret sauce." Instead of blasting the question to everyone, it targets Twitter users who have linked Foursquare accounts and have frequently checked into the relevant category/location.
- Forwarder: Once an answer is vetted, it is sent back to the original asker.

Experiments & Results: Crowds vs. Algorithms
The most striking result of the study is the comparison with Google. While Google is slightly better at factual queries (78% vs 75%), it collapses when faced with non-factual queries, dropping to a 29% answer rate. The crowdsourced system maintained a steady 75% success rate across both categories.
Performance Metrics
- Latency: You might think humans are slow, but 50% of answers arrived within 20 minutes.
- Expertise: Foursquare users provided significantly higher response rates for specific categories like Food and Nightlife compared to random local users.
- Mobile Dominance: 80% of answers came from mobile devices, proving that this is a "on-the-go" solution.
The data shows a clear correlation: higher quality questions (Rank 3) elicit higher quality, "Good Answers" from the crowd.
The Cumulative Distribution Function (CDF) shows that the system is viable for real-time needs, with 90% of answers arriving within 2 hours.
Critical Insight & Conclusion
The study proves that social context is a proxy for expertise. By utilizing "Check-ins," the researchers bypassed the need for complex profile analysis and went straight to physical proof of presence.
Takeaway: The future of search isn't just about better indexing; it's about better routing. As we move into an era of AI, this paper reminds us that the specialized, subjective knowledge of a "local mayor" on Foursquare is still a goldmine that generic models struggle to replicate.
Limitations: The system relies on the "goodwill" of strangers. Without a formal incentive or "karma" system, maintaining a 5% reply rate might be difficult as the system scales. However, the 9% retweet rate suggests that "social altruism" is a powerful, underutilized engine for information retrieval.
