Crowdsourcing the "Subjective" Map: Bridging the Gap in Location-Based Search

Crowdsourcing location-based queries

2011-03-01
Muhammed Fatih Bulut, Yavuz Selim Yilmaz, Murat Demirbas
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a framework for crowdsourcing location-based queries using Twitter as a middleware and Foursquare for expert identification. It focuses on non-factual, subjective queries where traditional search engines fail, achieving a 75% answer rate for such queries compared to Google's 29%.

TL;DR

Traditional search engines are great at finding "hotels in Miami," but they fail miserably when you ask for a "cheap, baby-friendly hotel near Broadway." This paper addresses this gap by crowdsourcing location-based queries to Foursquare experts via Twitter. The result? A system that answers 75% of subjective queries where Google manages less than 30%.

Problem & Motivation: The Failure of Factual Search

When we use a smartphone, our queries are often "non-factual." They are relative, multi-dimensional, and highly subjective. The authors discovered that 63% of location-based queries on Twitter fall into this category.

While algorithms can index static web pages, they cannot "feel" the vibe of a coffee shop or know if an apartment allows dogs based on a simple keyword search. The authors' insight is simple: Humans are the best sensors for subjective data. By identifying people who actually visit these locations (using Foursquare check-ins), we can route questions to "local experts" who are more likely to provide high-quality, nuanced answers.

Methodology: Routing Questions to the Right Crowd

The system architecture consists of a sophisticated pipeline designed to filter noise and ensure quality:

  1. Question Collector: Monitors Twitter for location-specific questions using a template-based approach (e.g., "Anyone... [location]... ?").
  2. Validator: A human-in-the-loop stage where moderators categorize the question (e.g., Food, Nightlife) and rank its quality.
  3. Asker (The Expert Finder): This is the "secret sauce." Instead of blasting the question to everyone, it targets Twitter users who have linked Foursquare accounts and have frequently checked into the relevant category/location.
  4. Forwarder: Once an answer is vetted, it is sent back to the original asker.

Overall System Architecture

Experiments & Results: Crowds vs. Algorithms

The most striking result of the study is the comparison with Google. While Google is slightly better at factual queries (78% vs 75%), it collapses when faced with non-factual queries, dropping to a 29% answer rate. The crowdsourced system maintained a steady 75% success rate across both categories.

Performance Metrics

  • Latency: You might think humans are slow, but 50% of answers arrived within 20 minutes.
  • Expertise: Foursquare users provided significantly higher response rates for specific categories like Food and Nightlife compared to random local users.
  • Mobile Dominance: 80% of answers came from mobile devices, proving that this is a "on-the-go" solution.

Question Ranks vs. Answer Ranks The data shows a clear correlation: higher quality questions (Rank 3) elicit higher quality, "Good Answers" from the crowd.

Response Time Distribution The Cumulative Distribution Function (CDF) shows that the system is viable for real-time needs, with 90% of answers arriving within 2 hours.

Critical Insight & Conclusion

The study proves that social context is a proxy for expertise. By utilizing "Check-ins," the researchers bypassed the need for complex profile analysis and went straight to physical proof of presence.

Takeaway: The future of search isn't just about better indexing; it's about better routing. As we move into an era of AI, this paper reminds us that the specialized, subjective knowledge of a "local mayor" on Foursquare is still a goldmine that generic models struggle to replicate.

Limitations: The system relies on the "goodwill" of strangers. Without a formal incentive or "karma" system, maintaining a 5% reply rate might be difficult as the system scales. However, the 9% retweet rate suggests that "social altruism" is a powerful, underutilized engine for information retrieval.

Find Similar Papers

Try Our Examples

  • Explore recent advancements in hybrid AI-crowdsourcing systems that specifically target subjective and non-factual query resolution in 2024-2025.
  • Which paper first introduced the "Aardvark" social search architecture, and how have subsequent works improved upon its routing algorithms via location-based services?
  • How can large language models (LLMs) be integrated with the Foursquare-Twitter crowdsourcing approach to pre-filter queries or synthesize multiple human answers?
Contents
Crowdsourcing the "Subjective" Map: Bridging the Gap in Location-Based Search
1. TL;DR
2. Problem & Motivation: The Failure of Factual Search
3. Methodology: Routing Questions to the Right Crowd
4. Experiments & Results: Crowds vs. Algorithms
4.1. Performance Metrics
5. Critical Insight & Conclusion