Social Media Mining: How Business Models and Privacy Settings Gatekeep Big Data
Social Media Mining: Impact of the Business Model and Privacy Settings
This paper presents a comparative case study of social media mining on Facebook and Twitter, focusing on how different business models and privacy configurations affect data accessibility. Using the 2015 Nepal earthquake as a benchmark, the study highlights the shift from open academic access to restricted, licensed data models.
Executive Summary
TL;DR: This paper investigates a fundamental but often ignored variable in Data Science: the platform itself. By comparing Facebook and Twitter during the 2015 Nepal earthquake crisis, the authors demonstrate that a platform’s business model (advertising vs. data licensing) and privacy settings are not just technical hurdles—they are filters that fundamentally bias the results of data mining.
Background: Positioned at the intersection of Information Systems and Data Mining, this work acts as a critical reality check for researchers, moving from "blindly mining data" to "critically assessing the source."
The Hidden Barriers: Why "Free" Data Isn't Free
The authors argue that social media platforms are no longer just communication hubs; they are profit-driven entities where users are the primary asset.
- Twitter's Strategy: Relies heavily on data licensing (the "Firehose"). While it offers a free Streaming API, it is restricted to a 1% sample, potentially skewing results for niche topics.
- Facebook's Strategy: A "walled garden" approach. Their pivot to Graph API v2.0+ effectively shut down automated mining for independent researchers to protect user privacy and proprietary data value.
The motivation for this study was to prove that these corporate decisions directly lead to Sampling Bias, making the "digital copy" of society provided by these platforms incomplete.
Methodology: A Tale of Two APIs
The researchers used the R statistical environment to pull data using specific search parameters related to the Nepal earthquake (e.g., #nepal, #earthquake).
The Technical Split
- Twitter (twitteR package): Automated, metadata-rich, but limited by rate-limiting windows.
- Facebook (Rfacebook & Manual): Automated access was blocked by API changes during the study. The authors were forced into manual collection, revealing that keyword searches on Facebook are heavily influenced by the researcher's own social graph—a massive blow to objectivity.

Experiments and Key Findings: Facebook vs. Twitter
The study compared the frequency and sentiment of words found in the collected datasets.
Quantifying the Impact
- Character Limits: Twitter’s 140-character limit (at the time) resulted in lower word frequencies per post but higher metadata density.
- The Keyword Trap: On Facebook, keyword searches returned results primarily from "connected pages," whereas hashtag searches were more "profile independent."
- Volume Disparity: The word "Nepal" appeared 1,242 times in the Facebook set compared to only 336 times on Twitter, showcasing how different platform architectures allow for more or less detailed text mining.
Figure: The dominance of specific disaster-related stems on Facebook demonstrates high information density but highlights the lack of automated filtering tools compared to Twitter.
Critical Analysis & Conclusion
Takeaway
The core contribution of this paper is the verification that Twitter is the superior platform for academic social media mining due to its API's support for automation and metadata, despite its 1% sampling limit. Facebook, while richer in personal sentiment, has effectively "legally and technically" locked its gates to the research community.
Limitations
- Temporal Scope: The study was conducted during a specific API transition period (2015), and the landscape has since become even more restrictive.
- Manual Bias: The forced manual collection on Facebook introduces human error that mirrors the very bias the authors intended to study.
Future Outlook
As major platforms (including Twitter/X) move toward paid API models, the "impact of the business model" discussed here will become the dominant factor in the feasibility of future social media research. The era of "free" academic access to the global digital conversation is rapidly closing.
