Crowdstore: Reimagining the Crowd as a Graph Database

Crowdstore: A Crowdsourcing Graph Database

2016-01-01
Vitaliy Liptchinsky, Benjamin Satzger, Stefan Schulte, Schahram Dustdar
Summary
Problem
Method
Results
Takeaways
Abstract

Crowdstore is a specialized hybrid graph database system that introduces CRowdPQ, an extension of Conjunctive Regular Path Queries. It enables complex, collaborative, and responsive crowdsourcing by incorporating human workers directly into the graph schema and query logic.

TL;DR

Crowdstore is a groundbreaking graph database management system that treats human workers not just as task-takers, but as nodes in a living data graph. By introducing CRowdPQ, the authors extend formal graph query languages to support synchronous collaboration, expert discovery, and real-time tasking, bridging the gap between machine precision and human intuition.

Background Positioning: This work moves beyond the "SQL-for-Crowds" era (like CrowdDB) into the collaborative graph era, focusing on the social and relational dynamics of human computation.

Problem & Motivation: The Limits of the "Pull" Model

Most hybrid databases treat crowdsourcing like a vending machine: you drop in a query (a HIT), and a random worker picks it up. This "pull model" has three fatal flaws:

  1. Isolation: Workers can't easily collaborate on complex, multi-skilled tasks (e.g., a designer and a developer working together).
  2. Subjectivity: Matching tasks to experts relies on workers' self-assessment rather than proven social or professional graphs.
  3. Latency: Recruiting a new crowd for every query is too slow for real-time applications.

The authors argue that because human relationships and task workflows are inherently relational, a Graph Database is the superior abstraction for hybrid computation.

Methodology: Extending Graph Theory to Humans

At the heart of the system is CRowdPQ, which builds on Conjunctive Regular Path Queries (CRPQ). It introduces two critical new types of relations:

  • Descriptor (DRPQ) <...>: These are free-text conditions interpreted by humans. For example, ?pic <"Draw a sheep"> ?worker.
  • Resolver (RRPQ) [...]: These define the dataflow—who produces the data and who consumes it. This allows for "denormalized joins" where workers see the context of the whole task.

System Architecture & The "Crowd Pool"

The genius of Crowdstore lies in how it adapts classic RDBMS concepts:

  1. Crowd Pool (The Buffer Manager): Just as a DB caches disk pages in RAM, Crowdstore keeps a set of workers "on-call" (on a small payroll) to ensure 2-3 second response times. It uses a Least-Recently-Active (LRA) algorithm to evict "cold" workers.
  2. Crowd Indexes: These are "Routing Indexes" where some workers (Index Workers) are paid specifically to find other expert workers, effectively crowdsourcing the search for talent.

Table of Comparison Table 1: Comparison of Crowdstore against existing hybrid databases.

Experiments & Expressiveness

The authors validated CrowdPQ through several high-complexity scenarios:

  • Social Vicinity: Finding two workers who are "friends of friends" to ensure they can collaborate effectively on a web design task.
  • Anti-Bias Filtering: Using a FILTER NOT EXISTS clause to ensure that the worker ranking a painting is not socially related to the artist.

Unlike previous systems, CrowdPQ allows the query writer to explicitly define Join Ordering via the consume relation. This is critical because human "joins" are expensive; knowing whether to filter by location first or by skill first can save thousands of dollars.

Critical Analysis & Conclusion

Takeaway

Crowdstore shifts the paradigm of crowdsourcing from "micro-tasking" to "social computation." By formalizing human interaction as graph paths, it enables the automation of complex workflows that were previously manual.

Limitations & Future Work

While the theoretical framework is robust, the system relies heavily on the quality of the social graph metadata. In practice, obtaining real-time social relations between transient crowd workers remains a challenge. Future iterations could integrate behavioral tracking to dynamically build these social graphs based on past collaborative success.

In conclusion, Crowdstore demonstrates that the "crowd" is not just a source of data, but a structured resource that can be indexed, cached, and queried with the same rigor as a traditional database.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply graph database optimization techniques to social crowdsourcing platforms or human computation workflows.
  • Which paper originally proposed the concept of "flash crowds" in crowdsourcing, and how does CrowdPQ’s formal grammar technically build upon that foundation?
  • Explore how the "Crowd Pool" buffer management logic could be extended using Reinforcement Learning to predict worker availability and optimize payroll budget.
Contents
Crowdstore: Reimagining the Crowd as a Graph Database
1. TL;DR
2. Problem & Motivation: The Limits of the "Pull" Model
3. Methodology: Extending Graph Theory to Humans
3.1. System Architecture & The "Crowd Pool"
4. Experiments & Expressiveness
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work