Smart Quizzing: Boosting Crowdsourced Knowledge Graphs via Worker Performance Prediction
Task Selection Based on Worker Performance Prediction in Gamified Crowdsourcing
This paper introduces a task selection method for gamified crowdsourcing designed to build a word retrieval assistant for people with aphasia. By transforming knowledge triple collection into four types of fill-in-the-blank quizzes and utilizing a performance prediction scoring function, the system matches specific quizzes to users most likely to know the answers.
TL;DR
Building a robust knowledge base often requires the collective intelligence of thousands of casual users. However, if you give a user a quiz they can't answer, they leave. This paper presents a gamified crowdsourcing framework that predicts which "fill-in-the-blank" quizzes a user can solve based on their history. In simulations, this predictive task selection doubled the efficiency of knowledge collection compared to traditional random assignment.
Background: Knowledge as a Game
Knowledge Graphs (KGs) are the backbone of modern AI, typically represented as triples (Subject, Predicate, Object). For specialized applications—like a word retrieval assistant for people with aphasia—generic world facts aren't enough. We need subjective, daily-life knowledge.
The authors leverage Gamification, turning data entry into a quiz game. But the friction point is the "SKIP" button: every time a user skips a question they don't know, the system loses momentum and efficiency.
The Core Challenge: The Distribution of Knowledge
Knowledge is not distributed equally. User A might know everything about "Fruits" but nothing about "Electronics." If a system blindly asks User A about laptop hardware, it wastes a session. The goal of this research is to minimize "SKIP" responses by matching the right quiz to the right user at the right time.
Methodology: Predictive Scoring Functions
The researchers developed four quiz types: fib_object, fib_predicate, fib_subject_object, and fib_predicate_object. To decide which one to serve, they compare two scoring strategies:
- Baseline (wo_pred): Simply selects tasks based on how much information is missing in the database (priority to the most "incomplete" subjects).
- Predictive (w_pred): Tracks a user's history (). If a user successfully answered a quiz about "Apples" () or the predicate "Color" () in the past, the system gives higher priority to new quizzes involving those same entities.
Mathematically:
The score for a quiz for user is calculated as: Where represents missing items and tracks historical success (+1) or failures/skips (-1).
Figure 1: The logic flow determining which type of quiz to generate based on current KB state.
Experiments & Results
The authors ran simulations comparing an "Omniscient User" (who knows everything) against 100 specialized users.
| Metric | Omniscient | Without Prediction | With Prediction |
|---|---|---|---|
| Quiz Efficiency (QE) | 0.806 | 0.054 | 0.129 |
| Total Quizzes Needed | 1,488 | 22,305 | 9,334 |
The results are striking: By simply tracking which subjects and predicates a worker was familiar with, the system reduced the total number of quizzes required to map 1,200 triples by over 58%.
Figure 2: The growth of the knowledge base over time. The "With Prediction" curve (blue) reaches completion much faster than the "Without Prediction" (orange).
Critical Insight: The "Skip" Problem
The breakdown of user responses revealed that fib_object quizzes suffered the most from "SKIP" responses. The predictive model succeeded primarily because it effectively "found" the users who possessed the missing objects for specific subject-predicate pairs.
Conclusion & Future Outlook
This work proves that even simple historical tracking can drastically improve the ROI of crowdsourcing. For developers of LLMs or Knowledge Graphs, the takeaway is clear: Don't treat your crowd as a monolith.
Future Work: The authors plan to integrate blockchain for data validation and explore more complex quiz types (like True/False verification) to further refine the accuracy of the collected knowledge.
Editor's Note: This research is a vital step toward creating "living" knowledge bases that evolve through casual human interaction, particularly for niche medical and social applications.
