Crowdstore: Reimagining the Crowd as a Graph Database
Crowdstore: A Crowdsourcing Graph Database
Crowdstore is a specialized hybrid graph database system that introduces CRowdPQ, an extension of Conjunctive Regular Path Queries. It enables complex, collaborative, and responsive crowdsourcing by incorporating human workers directly into the graph schema and query logic.
TL;DR
Crowdstore is a groundbreaking graph database management system that treats human workers not just as task-takers, but as nodes in a living data graph. By introducing CRowdPQ, the authors extend formal graph query languages to support synchronous collaboration, expert discovery, and real-time tasking, bridging the gap between machine precision and human intuition.
Background Positioning: This work moves beyond the "SQL-for-Crowds" era (like CrowdDB) into the collaborative graph era, focusing on the social and relational dynamics of human computation.
Problem & Motivation: The Limits of the "Pull" Model
Most hybrid databases treat crowdsourcing like a vending machine: you drop in a query (a HIT), and a random worker picks it up. This "pull model" has three fatal flaws:
- Isolation: Workers can't easily collaborate on complex, multi-skilled tasks (e.g., a designer and a developer working together).
- Subjectivity: Matching tasks to experts relies on workers' self-assessment rather than proven social or professional graphs.
- Latency: Recruiting a new crowd for every query is too slow for real-time applications.
The authors argue that because human relationships and task workflows are inherently relational, a Graph Database is the superior abstraction for hybrid computation.
Methodology: Extending Graph Theory to Humans
At the heart of the system is CRowdPQ, which builds on Conjunctive Regular Path Queries (CRPQ). It introduces two critical new types of relations:
- Descriptor (DRPQ)
<...>: These are free-text conditions interpreted by humans. For example,?pic <"Draw a sheep"> ?worker. - Resolver (RRPQ)
[...]: These define the dataflow—who produces the data and who consumes it. This allows for "denormalized joins" where workers see the context of the whole task.
System Architecture & The "Crowd Pool"
The genius of Crowdstore lies in how it adapts classic RDBMS concepts:
- Crowd Pool (The Buffer Manager): Just as a DB caches disk pages in RAM, Crowdstore keeps a set of workers "on-call" (on a small payroll) to ensure 2-3 second response times. It uses a Least-Recently-Active (LRA) algorithm to evict "cold" workers.
- Crowd Indexes: These are "Routing Indexes" where some workers (Index Workers) are paid specifically to find other expert workers, effectively crowdsourcing the search for talent.
Table 1: Comparison of Crowdstore against existing hybrid databases.
Experiments & Expressiveness
The authors validated CrowdPQ through several high-complexity scenarios:
- Social Vicinity: Finding two workers who are "friends of friends" to ensure they can collaborate effectively on a web design task.
- Anti-Bias Filtering: Using a
FILTER NOT EXISTSclause to ensure that the worker ranking a painting is not socially related to the artist.
Unlike previous systems, CrowdPQ allows the query writer to explicitly define Join Ordering via the consume relation. This is critical because human "joins" are expensive; knowing whether to filter by location first or by skill first can save thousands of dollars.
Critical Analysis & Conclusion
Takeaway
Crowdstore shifts the paradigm of crowdsourcing from "micro-tasking" to "social computation." By formalizing human interaction as graph paths, it enables the automation of complex workflows that were previously manual.
Limitations & Future Work
While the theoretical framework is robust, the system relies heavily on the quality of the social graph metadata. In practice, obtaining real-time social relations between transient crowd workers remains a challenge. Future iterations could integrate behavioral tracking to dynamically build these social graphs based on past collaborative success.
In conclusion, Crowdstore demonstrates that the "crowd" is not just a source of data, but a structured resource that can be indexed, cached, and queried with the same rigor as a traditional database.
