Hierarchical Skills: The Key to Optimizing Knowledge-Intensive Crowdsourcing
Using Hierarchical Skills for Optimized Task Assignment in Knowledge-Intensive Crowdsourcing
This paper introduces a hierarchical skill model for knowledge-intensive crowdsourcing, moving beyond flat tag-based systems. It proposes a taxonomy-based distance metric and two optimized task assignment heuristics, MatchParticipantFirst and ProfileHash, to improve result quality by matching tasks to the most suitable experts.
TL;DR
While platforms like Amazon Mechanical Turk excel at simple image labeling, "Knowledge-Intensive" tasks require specific expertise. This paper argues that the current "flat" keyword approach to skills is insufficient. By modeling skills as a taxonomy (tree) and using a specialized distance metric, the authors propose algorithms that match workers to tasks with surgical precision, significantly boosting output quality without sacrificing speed.
The Problem: The "Flat Skill" Limitation
In current crowdsourcing environments, skills are treated as isolated tags. If a task requires "Java 1.8" and a worker only lists "Java," the system might fail to see the connection. Conversely, it cannot easily infer that an expert in "Java 1.8 Threads" is highly qualified for a general "Java" task.
This lack of reasoning capability leads to:
- Low Precision: Keywords like "Java" might refer to the programming language or the Indonesian island.
- Inflexible Substitution: Systems cannot intelligently substitute a missing specific skill with a related, broader skill.
Methodology: Reasoning via Taxonomy
The authors propose a model where skills are nodes in a tree . The core innovation is the Skill Distance calculation:
- Insight: This formula favors common ancestors that are deeper in the tree. Two specialized sub-skills of "Machine Learning" are considered "closer" than two general skills under the root "Technology."
The ProfileHash Algorithm
To make this scalable for platforms with 500k+ users, the authors developed ProfileHash. It utilizes a series of hashmaps to index participant skills and their prefixes. When a task arrives, the system:
- Looks for an exact match.
- If none, it searches for more specialized workers (moving down the tree).
- If still none, it "relaxes" the requirement by moving up to a broader category (moving up the tree).
Figure 1: Example of a skill taxonomy used for reasoning.
Experimental Evidence
The authors compared their heuristics against the Hungarian Method (the gold standard for optimal matching) and ExactThenRandom (the keyword-based baseline).
1. Accuracy and Quality
In synthetic tests, the hierarchical approach consistently stayed closer to the optimal "Hungarian" lower bound than keyword matching. As the "skill budget" of participants increased, the taxonomy-based model was able to find much better matches than flat models.
Figure 2: Normalized cumulative distance (lower is better) vs. number of participants.
2. Real-World Validation (The Computer Science Quiz)
Using a real crowd from CrowdFlower and a laboratory group, participants were tested on a 58-question Computer Science quiz. The ProfileHash algorithm resulted in a higher ratio of correct answers across all groups (Lab, Crowdflower, and Mixed).
Figure 3: Accuracy improvement across different crowd types.
Critical Insights & Future Outlook
The beauty of this work lies in its Inductive Bias: it assumes that human knowledge is structured. By encoding this structure directly into the matching algorithm, we reduce the search space and improve the reliability of the crowd.
Limitations:
- The model currently assumes a single skill per task. Real-world complex tasks (e.g., "Full-stack development") often require a multi-skill intersection.
- It assumes the taxonomy is static and provided by experts, whereas professional skills evolve rapidly.
Takeaway: For any developer building a specialized talent platform or a collaborative science project, moving from "Tags" to "Taxonomies" is the single most effective way to improve the quality of your workforce assignments.
