EDBA: Precision Crowdsourcing for Web Accessibility Evaluation

A task assignment strategy for crowdsourcing-based web accessibility evaluation system

2017-04-02
Liangcheng Li, Can Wang, Shuyi Song, Zhi Yu, Fenqin Zhou, Jiajun Bu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Evaluator-Decision-Based Assignment (EDBA), a novel task assignment strategy for crowdsourcing-based web accessibility evaluation. It leverages a machine learning-based cost model and a greedy algorithm to match evaluation tasks with the most suitable volunteers, achieving State-of-the-Art accuracy in manual web accessibility audits.

TL;DR

Web accessibility evaluation often fails when automated tools meet complex, semantic checkpoints like CAPTCHA or Keyboard Traps. Crowdsourcing is the logical solution, but random task distribution leads to low accuracy. This paper presents EDBA (Evaluator-Decision-Based Assignment), a strategy that treats task assignment as an optimization problem. By modeling worker "costs" based on historical performance, it boosts evaluation accuracy and ensures a balanced workload for both novices and experts.

The Human Bottleneck in Accessibility

While automated tools catch basic HTML errors, they struggle with "contextualized" checkpoints. Can a blind user navigate this menu? Is the CAPTCHA truly perceivable? These require human judgment. However, professional experts are expensive and scarce.

Crowdsourcing platforms like Amazon Mechanical Turk offer scale but lack specialization. Assigning a "Keyboard Trap" check to a user with motor disabilities or a "CAPTCHA" check to a blind user (without proper tools) results in frustration and bad data. The core challenge is: How do we match the right task to the right volunteer without manual oversight?

Methodology: The EDBA Framework

The authors propose a two-stage approach to transform chaotic crowdsourcing into a precision engine.

1. Training the Cost Model

Instead of assuming all workers are equal, EDBA calculates a Cost for every evaluator-task pair. The cost is a weighted sum of three negative behaviors:

  • Error Rate (E): Inconsistency with expert reviews.
  • Give-Up Rate (G): Tasks abandoned by the user.
  • Time-Out Rate (T): Tasks that exceeded the time limit.

Using a Least Square Loss analysis and Gradient Descent, the system learns the optimal weights () for these factors by comparing them against an expert-assigned "Cost Assessment Value" (CAV).

2. The Greedy Optimization Algorithm

With the cost matrix ready, the system aims to minimize the total assignment cost: subject to constraints that ensure each task gets evaluators and no single worker is overwhelmed.

Assignment Procedure Figure 1: The EDBA workflow, from historical data ingestion to greedy task mapping.

Experimental Validation

The system was tested on the Chinese Web Accessibility Evaluation System using 20 distinct checkpoints (e.g., Multimedia, Error Suggestion, Link Position).

Accuracy Gains

The results showed that EDBA (red dots) consistently hit higher accuracy marks compared to the 1,000-run baseline of random assignment (blue boxplots). Even for difficult Level 3 checkpoints, EDBA provided a more reliable signal.

Performance Comparison Figure 2: Accuracy rate comparison across 20 checkpoint categories.

Ensuring "Fair" Participation

A critical finding was the Variance in Task Volume. Random assignment often leads to "worker burnout" or "newbie exclusion." EDBA maintained a variance of 0.95, whereas random assignment spiked to 5.62. This balance keeps volunteers engaged and prevents the system from relying too heavily on a few "super-users."

Critical Insight & Conclusion

The brilliance of EDBA lies in its dynamic nature. It doesn't just categorize a worker as "good" or "bad" once; it updates their Cost Matrix every few days. This allows novices to "warm up" on easier tasks and gradually move to high-weight checkpoints as their "Cost" decreases.

While the paper focuses on the Chinese standard (YD/T 1761-2012), the logic is universally applicable to any domain requiring Expertise-Aware Crowdsourcing, such as medical labeling or legal document review. The next frontier will likely involve determining the optimal number of evaluators () dynamically based on task difficulty.

Key Takeaways:

  • Behavior as Signal: Errors and timeouts are more than just bad data; they are features that define worker expertise.
  • Balance Matters: Reducing workload variance is as important for system sustainability as it is for accuracy.
  • EDBA is scalable: It bridges the gap between unreliable volunteers and expensive experts.

Find Similar Papers

Try Our Examples

  • Search for recent papers on expert-aware task assignment in knowledge-intensive crowdsourcing beyond web accessibility.
  • Which study first introduced the use of historical behavioral data to model worker reliability in crowdsourcing, and how does EDBA improve upon its loss function?
  • Explore how EDBA-style assignment strategies have been applied to multi-modal accessibility tasks involving AI-human collaboration.
Contents
EDBA: Precision Crowdsourcing for Web Accessibility Evaluation
1. TL;DR
2. The Human Bottleneck in Accessibility
3. Methodology: The EDBA Framework
3.1. 1. Training the Cost Model
3.2. 2. The Greedy Optimization Algorithm
4. Experimental Validation
4.1. Accuracy Gains
4.2. Ensuring "Fair" Participation
5. Critical Insight & Conclusion
5.1. Key Takeaways: