WinoFlexi: Scaling Commonsense Reasoning via Collaborative Crowdsourcing

WinoFlexi: A Crowdsourcing Platform for the Development of Winograd Schemas

2019-01-01
Nicos Isaak, Loizos Michael
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces WinoFlexi, the first dedicated crowdsourcing platform designed to scale the production of Winograd Schema Challenge (WSC) datasets. The system employs a collaborative "contributor-evaluator" model to generate linguistic pairs that test commonsense reasoning, achieving a quality level comparable to expert-crafted benchmarks like the Winograd-library.

TL;DR

The Winograd Schema Challenge (WSC) is a benchmark for machine intelligence that requires deep commonsense reasoning. Historically, WSC datasets have been tiny because they are hard for humans to write. This paper presents WinoFlexi, a platform that successfully utilizes crowdworkers to build high-quality Winograd schemas. By using a clever mix of automated validation and a "contributor-turned-evaluator" workflow, the authors proved that the "crowd" can match the quality of "experts" in creating complex linguistic riddles.

Background: The Problem of "Data Hunger" in WSC

Unlike standard NLP tasks, the Winograd Schema Challenge is specifically designed to be "Google-proof." It consists of pairs of sentences that differ by only one or two words, which flip the referent of a pronoun. For example:

  1. The trophy doesn't fit into the brown suitcase because it is too large. (it = trophy)
  2. The trophy doesn't fit into the brown suitcase because it is too small. (it = suitcase)

Since creating these requires high creativity and world knowledge, the field has relied on two tiny collections (Rahman & Ng and the Winograd-library) for years. WinoFlexi aims to break this bottleneck by turning crowdsourcing into a scalable "factory" for commonsense data.

Methodology: How WinoFlexi Ensures Quality

Crowdsourcing complex tasks often leads to "noise" or cheating. WinoFlexi mitigates this through a sophisticated architecture:

1. The Contributor-Evaluator Pipeline

The system doesn't just hire workers; it trains them.

  • Registration & Training: New users must resolve existing expert schemas to prove their English proficiency.
  • Contributor Role: Workers draft new schemas. The system performs real-time heuristic checks to ensure sentences, questions, and targets are related.
  • Evaluator Promotion: Only contributors who maintain an approval rate above 90% (matching adult human performance) are promoted to act as Evaluators for others' work.

2. Leakage and Hardness Monitoring

To prevent workers from simply copying known examples, WinoFlexi uses a Leakage Detector that compares new submissions against existing libraries. Furthermore, it incorporates an automated Hardness Metric to provide feedback to workers, encouraging them to write more challenging sentences if their drafts are too "easy" for machines to solve.

System Architecture Figure 1: The WinoFlexi platform architecture showing the flow from registration to payment verification.

Experimental Results: Crowd vs. Experts

The core question was: Is crowd-generated data actually good?

The authors tested the new schemas against three state-of-the-art coreference resolution systems (Stanford-Core-NLP, Wikisense, and K-Parser). The results were striking:

  • Performance Correlation: The accuracy of these AI systems on WinoFlexi schemas mirrored their performance on the expert-crafted Winograd-library with a correlation coefficient as high as 0.995.
  • Diversity: Crowdworkers explored subjects from "Spiderman vs Hulk" to "psychiatrists," showing higher creative variance than traditional academic datasets.

Hardness Comparison Figure 2: Violin plot showing the comparable hardness distribution between the expert Winograd-library and the WinoFlexi-library.

Critical Insight & Conclusion

The success of WinoFlexi lies in its Inductive Bias for quality control. By tying payment to a "Ban Score" and using high-performing peers as filters, the platform transforms a creative task into a structured pipeline.

Takeaway: The bottleneck for human-centric AI benchmarks isn't a lack of human creativity, but a lack of systems that can coordinate that creativity. WinoFlexi provides the blueprint for building larger, more diverse datasets that could finally push machines toward true commonsense reasoning.

Future Directions: The authors suggest "Human-Machine Teaming," where web-crawlers suggest potential sentence structures that crowdworkers then "Winograd-ize," potentially lowering the cost per schema even further.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Large Language Models (LLMs) to automatically generate or augment Winograd Schema Challenge datasets.
  • Which original paper by Terry Winograd or Hector Levesque defined the "selectional restrictions" pitfall in pronoun resolution, and how do modern benchmarks avoid it?
  • Explore research studies that compare the effectiveness of MicroWorkers versus Amazon Mechanical Turk for complex linguistic data annotation tasks.
Contents
WinoFlexi: Scaling Commonsense Reasoning via Collaborative Crowdsourcing
1. TL;DR
2. Background: The Problem of "Data Hunger" in WSC
3. Methodology: How WinoFlexi Ensures Quality
3.1. 1. The Contributor-Evaluator Pipeline
3.2. 2. Leakage and Hardness Monitoring
4. Experimental Results: Crowd vs. Experts
5. Critical Insight & Conclusion