Scaling Turkish PropBank: How Crowdsourcing Solves the Semantic Annotation Bottleneck
Verb Sense Annotation for Turkish PropBank via Crowdsourcing
This paper presents a methodology for constructing the Turkish PropBank verb sense corpus using crowdsourcing. By decomposing the complex Semantic Role Labeling (SRL) task into manageable microtasks, specifically Verb Sense Disambiguation (VSD), the author successfully annotated 5,855 verb instances with an 83.15% inter-annotator agreement using the CrowdFlower platform.
TL;DR
Building a PropBank (a corpus of semantically annotated predicates and arguments) is usually a slow, expert-heavy process. This paper presents a breakthrough for the Turkish PropBank by utilizing crowd intelligence to perform Verb Sense Disambiguation (VSD). By breaking the task into microtasks and enforcing strict quality controls, the author annotated 5,855 verb senses with 83.15% agreement in just 68 hours, proving that low-resource languages can scale their semantic resources affordably.
Motivation: The Semantic Data Gap
Semantic Role Labeling (SRL) is the bedrock of "shallow" semantic parsing—answering the question: Who did what to whom, when, and where? While languages like English and Chinese have mature PropBanks, Turkish faces two major hurdles:
- Morphological Complexity: A single Turkish word can contain multiple "Inflectional Groups" (IGs), making verb detection non-trivial.
- Resource Scarcity: Lack of expert annotators and funding for long-term manual labor.
The author's core insight is that while full SRL is too complex for laypeople, Verb Sense Disambiguation (choosing the right meaning from a list) is a task native speakers can perform with high accuracy if given the right interface and training.
Methodology: Engineering the Crowd
The researcher didn't just ask people to guess; they engineered a robust technical pipeline.
1. Preprocessing and Identification
Using the ITU-METU-Sabancı Treebank (IMST), the system identifies verbs even when they appear in derived forms (e.g., an adjective derived from a verb). To save costs, verbs with only one sense or those acting as copulas (like "ol" - to be) were filtered out or handled semi-automatically.
2. Job Design and Dynamic Rendering
Since different verbs have a different number of senses, the task interface was built using CML (CrowdFlower Markup Language). This allowed the UI to dynamically generate radio buttons for each sense found in the Turkish PropBank dictionary.
Figure 1: The annotation interface showing the predicate (Eylem) and the sentence (Cümle) context.
3. The Quality Control Trinity
The most critical part of the methodology is the three-step quality filter:
- Quiz Mode: Contributors must pass 5 test questions. 40% of applicants were rejected here.
- Gold Units: Hidden "test" questions were mixed into the real work. If a worker's accuracy dropped below 70%, they were automatically kicked out.
- Aggregation: Responses were weighted by "Contributor Trust," ensuring that the most reliable workers had more influence on the final label.
Results: Efficiency at Scale
The experiment yielded impressive throughput:
- Total Annotated: 5,855 senses.
- Speed: ~266 rows per hour.
- Cost: 0.05 per row).
- Reliability: Significant agreement (83.15%) and a very high confidence distribution among the final trusted judgements.
Figure 2: The distribution of judgements relative to contributor confidence, showing a heavy skew toward high-trust labels (above 0.7).
Critical Insight & Future Outlook
This paper is a significant "utility" contribution to Turkish NLP. It proves that the "Manual Labor" phase of linguistics can be bypassed using modern crowdsourcing platforms, provided the researchers pay close attention to Task Decomposition.
However, some challenges remain:
- "None" Options: About 700 rows were labeled as "None" by the crowd, indicating that the existing Turkish PropBank frame dictionary may still be incomplete and requires expert expansion.
- Next Step: The author plans to move to the much harder task of Argument Labeling (identifying who is the Agent, Patient, etc.), which will truly test the limits of non-expert crowd intelligence.
For practitioners in low-resource NLP, this work confirms that quality control is more important than initial worker expertise. By "coding" the instructions and quiz questions correctly, we can build high-quality SOTA datasets on a shoe-string budget.
