Mapping the Social Pulse: An Efficient Strategy for Facebook Place Data Extraction
An Efficient Facebook Place Information Extraction Strategy
The paper introduces a systematic strategy for large-scale data extraction from the Facebook Graph API using a hexagonal tiling approach. By bypassing common API limitations, the authors successfully mapped over 1.1 million places and nearly 959 million check-ins across Taiwan, specifically analyzing food-related trends like "Beef Noodle," "Sushi," and "Kimchi."
TL;DR
Researchers have developed a robust spatial extraction strategy to harvest massive datasets from Facebook's semi-closed ecosystem. By utilizing a hexagonal tiling method and automated API traversal, the study successfully mapped over 1.1 million locations in Taiwan, revealing deep insights into public behavior and culinary preferences, such as the overwhelming local popularity of Beef Noodles over international cuisines.
Background & Motivation: Moving Beyond Foursquare
In the world of Location-Based Social Networks (LBSN), most academic research has historically gravitated toward Foursquare or Twitter. Why? Because Facebook—the world’s largest social platform—is a "walled garden." Its API is notoriously restrictive and subject to frequent changes.
However, Facebook check-in data represents a goldmine for understanding Public Options—the collective behavior that defines "hot places" and high-density population movements. The authors identified a gap: while Taiwan has the highest Facebook penetration rate globally, scientific methods to extract its spatial data remained scarce. This paper aims to bridge that gap with a scalable, automated extraction framework.
Methodology: The Hexagonal Grid Strategy
The core technical challenge is that a single API request for a large area will often hit a "limit number" (n), failing to return all places in high-density urban centers.
1. Hexagonal Tiling (Circuit-Areas)
Instead of simple squares or random points, the authors use a hexagonal cell approach. They define a "unit area" where is the coordinate and is the radius. As seen in the architecture diagrams, they build layers of "Circuit-Areas" (CA):
- CA(1, C): The center cell.
- CA(2, C): The center plus 6 neighbors (7 total).
- CA(n, C): Expanding outward to perfectly cover a target administrative district.

2. Recursive Data Retrieval
For every cell, the system sends an HTTPS request to the Graph API. If the response contains a paging.next URL, the system recursively follows it until the entire cell's data is captured in JSON format and stored in a MySQL database.
3. Geographical Filtering
To ensure data purity, the authors apply an Odd-Even Rule point-in-polygon algorithm. This ensures that only places actually within the administrative boundary (like Xinyi District) are counted, even if the hexagonal grid slightly overlaps the border.
Experiments: The Culinary Map of Taiwan
The researchers demonstrated their system by scraping the entire island of Taiwan. The scale of the data is massive:
- Total Places: 1,112,188
- Total Check-ins: ~958.8 Million
- Population Average: ~40 check-ins per person.
Case Study: Beef Noodles vs. The World
To prove the system's utility for market research, the authors performed a keyword search across the extracted database for specific food types.
| Food Type | Number of Places | Percentage |
|---|---|---|
| Beef Noodle | 3,751 | 63% |
| Sushi | 1,978 | 33% |
| Kimchi | 252 | 4% |

The results provide a quantitative confirmation of Taiwan's national dish—Beef Noodles—while highlighting the smaller niche market for Korean cuisine (Kimchi) compared to Japanese (Sushi).
Critical Insight & Future Outlook
The primary value of this work lies in its data engineering logic. By treating geographic space as a searchable grid, the authors circumvent the "black box" limitations of the Facebook Search API.
Limitations: The study relies heavily on the "Name" field for categorization, which can be noisy (e.g., a "Sushi" place might also serve "Kimchi"). Future improvements could involve using Facebook’s category_list for multi-label classification.
Conclusion: This strategy transforms social media noise into structured "Big Data." For urban planners, it offers a way to detect population traffic; for businesses, it provides a precise tool for site selection and competitive analysis. The authors' future work aims to apply this to even more complex human behaviors, such as the impact of folk religions on social movement.
