Empowering the DHH Community: A Crowdsourced Approach to Real-Time Sign Language Interpretation

A Study Examining a Real-Time Sign Language-to-Text Interpretation System Using Crowdsourcing

2020-01-01
Kohei Tanaka, Daisuke Wakatsuki, Hiroki Minagawa
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a real-time sign language-to-text interpretation system that utilizes non-expert crowdsourcing, specifically involving Deaf and Hard-of-Hearing (DHH) individuals. The system segments live video into short, overlapping clips to distributed workers who provide text translations, achieving a collaborative captioning environment.

TL;DR

Bridging the gap between sign language and text has long been a resource-heavy challenge. This paper explores a novel crowdsourcing system that allows non-expert sign language readers (specifically DHH individuals) to provide real-time captions. By segmenting live video and providing specialized playback tools, the system aims to bypass the need for expensive professional interpreters while creating inclusion opportunities for the DHH community.

Background & Motivation: Moving Beyond the "Double Labor" Problem

In traditional settings, converting sign language to text is a "relay race":

  1. A professional interpreter watches the signer and speaks the translation.
  2. A captionist listens to that audio and types the text.

This workflow is not only expensive but inherently slow. Why not have sign language readers type the text directly? The primary barrier is cognitive load. Reading sign language is a visual task, and typing is a motor-intensive task; doing both simultaneously often leads to "sign oversight"—missing crucial information while looking at the keyboard or focused on the previous sentence.

Methodology: Slice, Translate, and Stitch

The research team developed a prototype based on Node.js and YouTube Live API to test a distributed workflow.

1. The Segmentation Strategy

The system divides the live stream into equal segments (set to 9 seconds in this study). To prevent loss of information at the boundaries, segments overlap by 1 second.

2. The Worker Interface (Task Execution)

To manage the cognitive demands, the interface includes:

  • Playback Control: 1-second rewind and pause.
  • Time Warp: Workers can slow down the video to type accurately and then use 1.5x speed to "catch up" to the live stream.
  • Neighbor Awareness: A "live typing" feed of the previous and next workers to help maintain context and avoid repeated words.

System Architecture Figure 1: The overall system architecture involving Task Control, Execution, and Display pages.

Experimental Analysis: The 3x Real-Time Barrier

The team conducted tests with four DHH university students. The results provided a sobering look at the challenges of non-expert crowdsourcing:

  • The Latency Gap: On average, it took a worker 26 seconds to interpret and type a 9-second clip. This implies that a single worker can never keep up with real-time sign language.
  • Typing Speed Bottleneck: The average typing speed was 1.9 Characters Per Second (CPS). Even the fastest worker (2.4 CPS) couldn't close the gap.
  • Context Loss: Because workers were so focused on their own segments, only one worker actually looked at the "neighboring" text. This resulted in a 66% rate of either missing text or colliding (overlapping) text between segments.

Performance Data Table Table 1: Worker typing speed (CPS) vs. task completion time (ms).

Critical Insight: The "15-Worker" Rule

One of the most valuable contributions of this study is the mathematical estimation of the workforce needed. By calculating the ratio of task time () plus "catch-up" time () over the segment length (), the authors concluded that 15 workers are necessary to handle one live stream sustainably.

Interestingly, the study found that workers who were assigned fewer tasks reported higher "enjoyment." This suggests that "gamifying" the experience or rotating workers frequently is key to preventing burnout in volunteer-based or crowdsourced platforms.

Conclusion & Future Outlook

While the current prototype faces challenges with text collisions and high latency, it proves a vital concept: DHH individuals can transition from being "consumers" of accessibility to "providers" of it.

The authors suggest that future iterations should include:

  • Dynamic Task Control: Adjusting segment lengths based on individual working memory and typing speed.
  • Correction Roles: Introducing a "Reviewer" tier in the crowd to stitch together the 9-second segments and fix the 66% collision/omission error rate.

This research is a stepping stone toward a more decentralized and inclusive digital world where the community supports itself through technology.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize multi-user collaborative editing to reduce latency in real-time sign language translation.
  • Which paper first proposed the "Scribe" or "TimeWarp" mechanisms for crowdsourced captioning, and how do they differ from visual-to-text translation requirements?
  • Investigate the use of State Space Models (SSM) or automated sign language recognition (SLR) to assist human workers in crowdsourced interpretation tasks.
Contents
Empowering the DHH Community: A Crowdsourced Approach to Real-Time Sign Language Interpretation
1. TL;DR
2. Background & Motivation: Moving Beyond the "Double Labor" Problem
3. Methodology: Slice, Translate, and Stitch
3.1. 1. The Segmentation Strategy
3.2. 2. The Worker Interface (Task Execution)
4. Experimental Analysis: The 3x Real-Time Barrier
5. Critical Insight: The "15-Worker" Rule
6. Conclusion & Future Outlook