Crowdsourcing in Computing: Mapping a Decade of Empirical Evidence
A State-of-the-Art of Empirical Literature of Crowdsourcing in Computing
This paper presents a Systematic Mapping Study (SMS) of empirical research in crowdsourcing within the computing and software engineering domains. By synthesizing 400 primary studies published between 2005 and 2015, the authors identify key trends in research methodologies, output types, and the maturation of the field from its inception to a more established state.
TL;DR
This research provides a comprehensive Systematic Mapping Study (SMS) of 400 empirical papers on crowdsourcing in the computing field. It reveals a field that has rapidly matured since 2010, dominated by experimental methodologies and the use of Amazon Mechanical Turk, yet reveals a critical gap: while we are good at using the "crowd" for testing and data annotation, we have yet to fully lean on them for core software design and coding.
Contextualizing the "Wisdom of the Crowd"
Crowdsourcing—a term coined by Jeff Howe in 2006—has fundamentally altered the Inductive Bias of software production. We have moved from "cathedral-style" collocated teams to a "bazaar-style" global call. However, for a long time, the academic community struggled to quantify how this paradigm was actually being researched. Was it just a buzzword, or was there rigorous empirical backing?
The "State of the Art" Gap
The authors highlight a significant maturation of the field within a remarkably short window (less than a decade). However, they identify a "surface-level" maturity. The problem is twofold:
- Methodological Narrowness: A staggering 81% of studies are experiments, but many lack rigorous experimental design.
- Domain Imbalance: Most empirical work focuses on crowdsourced testing or simple data annotation (image/text tagging), leaving the complex "creative" parts of software engineering—like requirements engineering and architectural design—largely untouched.
Methodology: How the Mapping was Conducted
The researchers followed established evidence-based software engineering (EBSE) guidelines to filter thousands of papers down to 400 primary studies. They analyzed these across several dimensions:
- Temporal Distribution: Tracking the explosion of interest from 2008 to 2015.
- Output Types: Categorizing results into guidelines, tools, techniques, and models.
- Platform Usage: Identifying which "digital labor markets" actually power the research.
Key Findings & Visual Evidence
Through their analysis, several striking trends emerged:
1. The Dominance of Experimentation
Unlike other areas of SE where case studies and action research are common, Crowdsourcing is heavily experimental. However, the authors critique the quality of these experiments, noting a lack of proper design description and a complete absence of experience reports.
2. Platform Hegemony
Amazon Mechanical Turk (AMT) is the undisputed king of crowdsourcing research, used in 120 studies. While new platforms are proposed frequently (40+ new platforms were identified in the literature), very few of them are empirically evaluated for their strengths or weaknesses.
3. The SE Lifecycle Gap
As shown in the paper's distribution analysis, the "Software Testing" phase has successfully adopted crowdsourcing, likely due to the parallelizable nature of bug hunting. Contrastingly, the coding and design phases are still in their infancy.
Critical Insight: Is More Always Better?
The paper concludes with a sober warning. While the quantity of studies suggests a mature field, the nature of the research suggests we are still scratching the surface.
- No Replication: There is almost no replication of existing studies, which is a cornerstone of the scientific method.
- Coordination Challenges: The industry still struggles with the "human" overhead of crowdsourcing—communication, task allocation, and shared understanding in a global context.
Takeaway for Future Researchers
This study serves as a roadmap. If you are looking for a PhD topic or a research gap, look toward Crowdsourced Requirements Engineering or Design. The community has enough guidelines and tools; what we need now is an evaluation of why those tools work and how to manage the complex human coordination required for high-level software tasks.
Note: This blog is based on the SMS results from 2005-2015. Since then, the advent of AI-assisted coding may have shifted these dynamics even further—a fascinating area for a follow-up study.
