EDBA: Precision Crowdsourcing for Web Accessibility Evaluation
A task assignment strategy for crowdsourcing-based web accessibility evaluation system
The paper introduces Evaluator-Decision-Based Assignment (EDBA), a novel task assignment strategy for crowdsourcing-based web accessibility evaluation. It leverages a machine learning-based cost model and a greedy algorithm to match evaluation tasks with the most suitable volunteers, achieving State-of-the-Art accuracy in manual web accessibility audits.
TL;DR
Web accessibility evaluation often fails when automated tools meet complex, semantic checkpoints like CAPTCHA or Keyboard Traps. Crowdsourcing is the logical solution, but random task distribution leads to low accuracy. This paper presents EDBA (Evaluator-Decision-Based Assignment), a strategy that treats task assignment as an optimization problem. By modeling worker "costs" based on historical performance, it boosts evaluation accuracy and ensures a balanced workload for both novices and experts.
The Human Bottleneck in Accessibility
While automated tools catch basic HTML errors, they struggle with "contextualized" checkpoints. Can a blind user navigate this menu? Is the CAPTCHA truly perceivable? These require human judgment. However, professional experts are expensive and scarce.
Crowdsourcing platforms like Amazon Mechanical Turk offer scale but lack specialization. Assigning a "Keyboard Trap" check to a user with motor disabilities or a "CAPTCHA" check to a blind user (without proper tools) results in frustration and bad data. The core challenge is: How do we match the right task to the right volunteer without manual oversight?
Methodology: The EDBA Framework
The authors propose a two-stage approach to transform chaotic crowdsourcing into a precision engine.
1. Training the Cost Model
Instead of assuming all workers are equal, EDBA calculates a Cost for every evaluator-task pair. The cost is a weighted sum of three negative behaviors:
- Error Rate (E): Inconsistency with expert reviews.
- Give-Up Rate (G): Tasks abandoned by the user.
- Time-Out Rate (T): Tasks that exceeded the time limit.
Using a Least Square Loss analysis and Gradient Descent, the system learns the optimal weights () for these factors by comparing them against an expert-assigned "Cost Assessment Value" (CAV).
2. The Greedy Optimization Algorithm
With the cost matrix ready, the system aims to minimize the total assignment cost: subject to constraints that ensure each task gets evaluators and no single worker is overwhelmed.
Figure 1: The EDBA workflow, from historical data ingestion to greedy task mapping.
Experimental Validation
The system was tested on the Chinese Web Accessibility Evaluation System using 20 distinct checkpoints (e.g., Multimedia, Error Suggestion, Link Position).
Accuracy Gains
The results showed that EDBA (red dots) consistently hit higher accuracy marks compared to the 1,000-run baseline of random assignment (blue boxplots). Even for difficult Level 3 checkpoints, EDBA provided a more reliable signal.
Figure 2: Accuracy rate comparison across 20 checkpoint categories.
Ensuring "Fair" Participation
A critical finding was the Variance in Task Volume. Random assignment often leads to "worker burnout" or "newbie exclusion." EDBA maintained a variance of 0.95, whereas random assignment spiked to 5.62. This balance keeps volunteers engaged and prevents the system from relying too heavily on a few "super-users."
Critical Insight & Conclusion
The brilliance of EDBA lies in its dynamic nature. It doesn't just categorize a worker as "good" or "bad" once; it updates their Cost Matrix every few days. This allows novices to "warm up" on easier tasks and gradually move to high-weight checkpoints as their "Cost" decreases.
While the paper focuses on the Chinese standard (YD/T 1761-2012), the logic is universally applicable to any domain requiring Expertise-Aware Crowdsourcing, such as medical labeling or legal document review. The next frontier will likely involve determining the optimal number of evaluators () dynamically based on task difficulty.
Key Takeaways:
- Behavior as Signal: Errors and timeouts are more than just bad data; they are features that define worker expertise.
- Balance Matters: Reducing workload variance is as important for system sustainability as it is for accuracy.
- EDBA is scalable: It bridges the gap between unreliable volunteers and expensive experts.
