Beyond Algorithms: Transforming Social Networks into Human-Powered Data Engines
Web Data Management through Crowdsourcing Upon Social Networks
This paper introduces a framework for Web data management that leverages crowdsourcing within existing social networks (Facebook, LinkedIn, Twitter). It formalizes a task model (Input, Task Type, Output, Execution) to transform human insights into structured data updates.
TL;DR
The explosion of Web data has outpaced the ability of automated systems to judge quality, sentiment, and nuance. This paper proposes a formal framework to bridge traditional Web systems with social networks, treating "the crowd" as a programmable data source. By mapping data management tasks (like ranking or tagging) to social interactions (like 'Likes' or 'Comments'), the authors demonstrate that human insight can be structured and integrated directly into database schemas.
The "Human Insight" Gap
While search engines are excellent at retrieving factual data, they fail at tasks requiring subjective judgment—what the authors call "sense extraction." Traditional crowdsourcing (e.g., Mechanical Turk) fills this gap but treats humans as isolated, anonymous "processors."
The authors argue that the Social Context is the missing ingredient. We trust a recommendation from a colleague more than a random worker. The challenge? How do we formalize "asking a friend" into a repeatable, scalable data management workflow?
Methodology: The Anatomy of a Crowd Task
The paper introduces a formal tuple for a Crowdsourcing Data Management Task: <I, D, T, O, E>.
1. The Task Model
- Input (I): The raw data (Relational, Semantic, or Unstructured).
- Task Type (T): Operations categorized into a taxonomy (Preference, Annotate, Rank, Cluster, Describe, Update).
- Output (O): The modified data. Crucially, tasks are either Schema Modifiers (adding new columns like
score) or Tuple Modifiers (changing existing values).

2. Execution & Deployment
The system interfaces with platforms like Facebook and Twitter through two main strategies:
- Native API Access: Using built-in features (e.g., Facebook Likes) to collect data, which ensures higher user engagement but offers less flexibility.
- Embedded Applications: Creating custom interfaces within the platform for more complex logic.
Taxonomy of Human Actions
The paper provides a comprehensive mapping of human social actions to database operations.

For example, a "Like" task is modeled as a counter increment in the database schema, while a "Group" task acts as a clustering algorithm where humans define the boundaries between data points.
Experimental Insights: The Diffusion Effect
The prototype was tested on an "Investment/Job Hunting" scenario. One of the most fascinating findings was the Diffusion Effect. Unlike traditional crowdsourcing, where you get exactly what you pay for (or less), social crowdsourcing tasks often received more answers than invitations sent. This is due to the inherent "virality" of social networks, where friends of friends contribute to the task autonomously.

Critical Analysis & Future Outlook
Contribution: The work successfully formalizes "informal" social interaction into a structured "data manipulation" language. It moves crowdsourcing away from the "task market" model toward a "social community" model.
Limitations: The reliance on platform-specific APIs makes the system vulnerable to changes in social media policies (e.g., API restrictions that have become much stricter since the paper's publication). Reliability is also an issue; unlike paid workers, social "friends" have no contractual obligation to be accurate or timely.
The Future: As AI continues to generate massive amounts of content, the need for Human-in-the-loop verification in social contexts will only grow. This framework provides the groundwork for how we might "program" social communities to help AI stay grounded in human values and subjective truths.
