SimilarHITs: Decoding the DNA of Task Similarity in Crowdsourcing

SimilarHITs: Revealing the Role of Task Similarity in Microtask Crowdsourcing

2018-07-03
Alan Aipe, Ujwal Gadiraju, U. Gadiraju
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces SimilarHITs, a framework investigating task similarity in microtask crowdsourcing. It identifies 12 similarity dimensions and proposes a supervised machine learning model using an SGD Regressor to predict overall task similarity, exploring its impact on worker performance and engagement.

TL;DR

In the world of microtask crowdsourcing (like Amazon Mechanical Turk), the order in which workers complete tasks—known as task chaining—is vital. This paper reveals that while keeping tasks similar boosts accuracy and speed, it also leads to faster boredom. The sweet spot for worker retention actually lies in Random chains that provide enough variety to keep the mind engaged while maintaining some continuity.

Problem & Motivation: The Cost of Context Switching

Most crowdsourcing platforms treat Human Intelligence Tasks (HITs) as isolated atoms. In reality, workers consume them in chains. Previous research focused on task complexity, but ignored similarity.

The authors suggest that jumping between wildly different tasks (e.g., transcribing audio then identifying punctuation errors) forces a "re-configuration" of psychological parameters. This mental overhead results in slower completion times and higher error rates. Conversely, too much similarity might lead to monotonous "zombie-mode" work.

Methodology: How Do We Define "Similar"?

The authors identified 12 dimensions of similarity, discovering that Task Type and Workflow were the most influential in how workers perceive similarity.

The Model

To quantify similarity, they used an SGD Regressor to predict values for complex dimensions like "Topic" and "Workflow," combining them into an overall score via a weighted mean.

Model Parameters and Features

Table: The parameters used to train the SGD Regressors for Workflow and Topic similarity.

Experiments: The Triple-Chain Test

The study compared three conditions:

  1. Similar: High similarity between consecutive tasks.
  2. Dissimilar: Low similarity.
  3. Random: A baseline of mixed tasks.

Key Findings

  • Accuracy: Similar chains are the gold standard for quality. Workers in the Similar group achieved ~80% accuracy, significantly outperforming the Dissimilar group (~61%).
  • Retention & Boredom: This is where the "Similarity Trap" appears. While similar tasks were done well, they were also the most boring. Interestingly, the Random condition saw the highest retention, likely because the variety acted as "micro-diversions" that refreshed the worker's attention.

Worker Accuracy Over Tasks

Figure: Accuracy remains highest in the Similar condition even as workers progress through the chain.

Critical Analysis & Conclusion

Takeaway

The paper successfully proves that task similarity is a double-edged sword. If a requester needs high-precision data (e.g., medical labeling), they should prioritize similar task chains. If they need high volume and long-term engagement, they should introduce strategic "dissimilar" tasks to break the monotony.

Limitations

The study was conducted on CrowdFlower (now part of Appen) and relied on a specific set of 61 AMT tasks. The "Random" condition's success might be highly dependent on the quality of the mix, which wasn't deeply explored—i.e., how much variety is "just right"?

Future Outlook

This work paves the way for Dynamic Task Allocation. Imagine an AI-driven platform that monitors a worker's speed; when it detects slowing (boredom), it injects a dissimilar "palate cleanser" task to re-engage them, and subsequently switches back to similar tasks to maintain high accuracy.

Find Similar Papers

Try Our Examples

  • Find recent papers that explore the "balance between cognitive variety and task consistency" to optimize worker performance in crowdsourcing marketplaces.
  • Which study first introduced the "goal-oriented taxonomy of microtasks," and how has it been used to measure similarity in subsequent human-computer interaction research?
  • Search for research applying task similarity and chaining optimization to complex AI-assisted workflows like RLHF (Reinforcement Learning from Human Feedback) or multi-modal data labeling.
Contents
SimilarHITs: Decoding the DNA of Task Similarity in Crowdsourcing
1. TL;DR
2. Problem & Motivation: The Cost of Context Switching
3. Methodology: How Do We Define "Similar"?
3.1. The Model
4. Experiments: The Triple-Chain Test
4.1. Key Findings
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook