KT4Crowd: Predicting Expert Performance by Tracing Worker Knowledge Evolution

Predicting Crowdsourcing Worker Performance with Knowledge Tracing

2020-01-01
Zizhe Wang, Hailong Sun, Tao Han
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces KT4Crowd, a novel framework that adapts Knowledge Tracing (KT) models—originally designed for Intelligent Tutoring Systems—to predict worker performance in Knowledge-Intensive Crowdsourcing (CKI-C). By utilizing the DKVMN model within this framework, the authors achieve State-of-the-Art (SOTA) results in predicting developer success on platforms like Topcoder.

TL;DR

Predicting who will win a high-stakes software competition is notoriously difficult because expert skills are not static. This paper introduces KT4Crowd, a framework that treats competitive crowdsourcing like an Intelligent Tutoring System. By breaking down complex tasks and applying Knowledge Tracing (KT) models (specifically DKVMN), the authors can track how a worker's "knowledge state" evolves over time, leading to performance predictions that significantly outperform traditional Elo-style rating systems.

Background: The Problem with Static Ratings

In competitive platforms like Topcoder or Kaggle, workers aren't just labels; they are learners. Traditional worker models often assume a fixed "ability score" (like a chess rating). However, in Knowledge-Intensive Crowdsourcing (KI-C):

  1. Skills Evolve: A developer's proficiency in "Java" or "Machine Learning" changes with every task they complete.
  2. Multi-Dimensional Tasks: A single task might require UI design, algorithm optimization, and database management simultaneously.
  3. No Standard Answer: Unlike a math quiz, performance is measured in relative ranks and continuous scores.

The Insight: Crowdsourcing as a "Classroom"

The authors realized that CKI-C mirrors Intelligent Tutoring Systems (ITS). In both cases, we want to know: If a user has done X and Y in the past, can they solve Z now? To bridge the gap between educational KT and complex crowdsourcing, the KT4Crowd framework introduces two critical adaptations.

Methodology: The KT4Crowd Framework

The core innovation lies in how data is "translated" for the AI:

  • Task Decomposition: Since KT models (like DKT or DKVMN) usually handle one concept at a time, KT4Crowd breaks a task with features into separate subtasks.
  • Results Transformation: Instead of raw scores, the model focuses on "Good Performance"—defined as achieving a rank and score above specific thresholds.

Architecture of the Prediction Process

The framework utilizes a trained KT model to project a worker's future performance. It simulates a "loss" (0) for each subtask, passes it through a Memory Network, and uses the resulting state to predict the probability of success.

KT4Crowd Prediction Flow Figure 1: The prediction workflow showing the transformation of historical sequences into subtask predictions followed by Majority Voting.

Experiments and Results

The authors tested their framework against traditional rating systems (Glicko-2) and standard KT applications on a massive dataset of 50,625 submissions from Topcoder.

Sequence Prediction Performance

The comparison involved DKT (Deep Knowledge Tracing) and DKVMN (Dynamic Key-Value Memory Networks).

MetricDKT-S (Ours)DKVMN-S (Ours)DKT-O (Baseline)
AUC0.74670.84120.7166
Accuracy0.71250.82540.6980
F1 Score0.48070.75400.5498

Key Takeaways from the Data:

  • DKVMN-S is the clear winner: Utilizing a memory-augmented neural network allows the model to store specific skill proficiencies more effectively than a standard LSTM (DKT).
  • The Framework Matters: The "-S" (KT4Crowd) versions consistently outperformed the "-O" (One-feature) versions, proving that decomposing tasks into multiple skills is essential for accuracy.

Winner Prediction: Beating the Industry Standard

When it came to predicting specifically who would "win" or perform excellently on a new task, KT4Crowd (with DKVMN) outperformed the actual Topcoder Rating System and the Glicko-2 system by a wide margin (AUC 0.78 vs traditional methods failing to capture the complexity).

Performance Metrics Table Table: Comparison of KT4Crowd against Glicko and Topcoder rating systems.

Critical Insight & Conclusion

Why does this work? Traditional rating systems are summaries; Knowledge Tracing is a history. By maintaining a memory of how a worker handled specific skills in the past, the model can navigate the "Expertise Cold Start" problem and the "Skill Decay" problem simultaneously.

Limitations: The model currently relies on "Majority Voting" for subtasks, which assumes all skills in a task are equally important. Future iterations might benefit from an Attention Mechanism to weight the "Key Feature" of a task more heavily than secondary requirements.

Ultimately, KT4Crowd proves that the boundary between "learning" and "working" is porous. If we can trace how knowledge is acquired, we can predict how it will be applied in the global digital economy.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply Deep Knowledge Tracing (DKT) or Dynamic Key-Value Memory Networks (DKVMN) to professional skill assessment outside of traditional classroom settings.
  • What are the foundational papers for the Glicko-2 and Topcoder rating systems, and how do they mathematically differ from the temporal dependency modeling found in LSTM-based Deep Knowledge Tracing?
  • Explore how Task Decomposition methods in crowdsourcing are being integrated with Multi-Task Learning (MTL) to handle inter-dependencies between different worker skills.
Contents
KT4Crowd: Predicting Expert Performance by Tracing Worker Knowledge Evolution
1. TL;DR
2. Background: The Problem with Static Ratings
3. The Insight: Crowdsourcing as a "Classroom"
4. Methodology: The KT4Crowd Framework
4.1. Architecture of the Prediction Process
5. Experiments and Results
5.1. Sequence Prediction Performance
5.2. Winner Prediction: Beating the Industry Standard
6. Critical Insight & Conclusion