CDT-Framework: Mastering the Complexity of Multi-phase Crowdsourcing Tasks
16762_A Decision Tree Based Quality Control Framework for Multi-phase Tasks in Crowdsourcing.
This paper introduces a quality control framework for Multi-phase Tasks (MPTs) in crowdsourcing, such as travel planning or micro-writing. It utilizes a Constrained Decision Tree (CDT) and a probabilistic graphical model to manage task generation and result inference, achieving SOTA results in user satisfaction and cost efficiency.
TL;DR
Crowdsourcing complex tasks like travel planning or translation often fails because systems either ignore the logic between steps or become prohibitively expensive. This paper introduces a Constrained Decision Tree (CDT) framework that treats multi-phase tasks as a sequence of dependent decisions. By modeling worker reliability and task constraints (like travel distance) mathematically, the authors achieved a 63.3% boost in result quality while significantly slashing costs.
The "Broken Link" in Complex Crowdsourcing
Current crowdsourcing platforms excel at simple labels (e.g., "Is this a cat?"). However, they struggle with Multi-phase Tasks (MPTs).
Consider travel planning:
- The "Unified" Failure: If you ask a worker to plan a whole 10-stop trip, they might focus only on their niche interests, missing popular spots.
- The "Split" Failure: If you ask different workers to pick "Top-K" spots separately, they might pick famous locations that are 5 hours apart, making the route physically impossible.
The missing ingredient is the Constrained Relationship—the logical and physical threads that bind one subtask to the next.
Methodology: The Constrained Decision Tree (CDT)
The authors transform the workflow into a tree-search problem. Each level of the tree represents a "phase" of the task.
1. Probabilistic Result Inference
To handle noisy worker input, the framework uses a Probabilistic Graphical Model. It doesn't just "count votes"; it calculates the Improvement Ratio. It models:
- : Task difficulty based on how similar two options are.
- : Worker accuracy, which naturally drops as the difficulty increases.

2. The CDT Model
The core innovation is the Synthetical Score, which balances two factors:
- Satisfaction Score: How much the crowd likes the sequence.
- Transition Score: How well the sequence obeys constraints (e.g., traffic congestion, distance).
This formula ensures the final output isn't just "popular" but also "practical."

Experiments: More Quality, Less Spend
The researchers tested their system, CrowdTP, against the existing state-of-the-art, CrowdPlanr, using real-world Beijing travel scenarios.
1. Slashing Crowdsourcing Costs
By using a Top-K greedy search on the decision tree, the system avoids generating thousands of unnecessary sub-tasks. The data shows that CrowdTP required roughly 19,200 fewer subtasks than the baseline to reach an optimal plan.

2. Higher User Satisfaction
When independent judges evaluated the generated travel routes, CrowdTP consistently won. In one specific task, the satisfaction ratio was 63.3% higher than the baseline, proving that incorporating traffic and transit constraints is vital for "real" utility.
Critical Insight & Conclusion
This work highlights that Crowdsourcing is not just about data aggregation; it's about workflow orchestration.
The CDT framework succeeds because it acknowledges that human workers have limits (error rates) and that tasks have physics (constraints). While the greedy Top-K approach might technically risk local optima, the experimental evidence suggests it provides a massive efficiency gain for complex, sequential human-computing tasks. As we move toward more complex AI-Human hybrids, these "constrained" structures will be essential to keeping costs manageable.
