ICS Caption Editor: Leveraging Crowdsourcing to Bridge the ASR Accuracy Gap in STEM Education

A crowdsourcing caption editor for educational videos

2014-10-01
Rucha Deshpande, Tayfun Tuna, Jaspal Subhlok, Lecia Barker
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents the ICS Caption Editor, a web-based crowdsourcing framework designed to generate and refine captions for STEM educational videos. By integrating Automatic Speech Recognition (ASR) with a collaborative student-driven editing workflow, the system achieves near-perfect caption accuracy (99%) for complex technical lectures.

TL;DR

The ICS Caption Editor is a collaborative platform that turns the tedious task of lecture captioning into a manageable crowdsourced activity. By combining imperfect Automatic Speech Recognition (ASR) with student-led micro-tasks, the system achieves 99% accuracy, making technical STEM videos searchable and highly accessible without the high cost of professional transcription.

The "Broken" State of Automated Transcription

While LLMs and modern AI have improved speech-to-text, this paper highlights a persistent struggle in academia: Technical Context. In live STEM lectures, the combination of complex terminology (e.g., "Ungrammatical constructs" due to math variables), classroom echoes, and diverse instructor accents creates a "perfect storm" of errors.

The authors' pre-study found that even commercial tools like Dragon Naturally Speaking and YouTube ASR averaged only 68% to 85% accuracy for live lectures. This makes the resulting captions more of a distraction than a help. Manual professional services are the gold standard but are far too expensive for every-day classroom use.

Methodology: The Power of the Crowd

The core insight of the ICS Caption Editor is that students are the best domain experts for their own courses. If 10-12 students spend 45 minutes each, an 80-minute lecture is captioned perfectly in a few days.

Key Architectural Features:

  1. Micro-Partitioning: The system breaks the transcript into segments of 5 sentences. This prevents "task fatigue" and allows multiple students to work simultaneously.
  2. Audio Looping & Variable Speed: To catch difficult phrases, the editor loops the specific audio segment and provides a "PlaySpeed" tool to slow down playback without pitch distortion.
  3. The "Review" Protocol: Users can mark segments as "Needs Review," creating a multi-pass verification system that ensures technical terms are triple-checked.

ICS Caption Editor Interface Figure 1: The ICS Caption Editor interface featuring synchronized video, text editing, and status tracking (Complete vs. Needs Review).

Experimental Results: Accuracy Meets Efficiency

The evaluation focused on two Computer Science courses. The results were striking:

  • Near-Perfect Accuracy: Final captions reached 99% accuracy, far surpassing any ASR-only approach.
  • Distributed Effort: While one person would take ~10 hours to caption a lecture, the crowdsourced group completed it with a median individual effort of just 45 minutes.
  • Subjective Value: An overwhelming majority of students (most of whom were non-native English speakers) reported that captions improved their Efficiency, Note-taking, and Learning (see Figure 2).

Student Learning Impact Figure 2: Survey data showing that students perceived significant improvements in learning and attention due to the presence of captions.

Academic Insight: Why it Works

The success of the ICS Editor is rooted in collaborative validation. In technical domains, "Tool Weakness" (incorrect ASR hypothesis) accounts for 50% of errors. By providing students with a visual frame of the lecture slides alongside the audio, they use visual cues to correct what the ASR "heard" incorrectly.

Critical Analysis & Conclusion

Takeaway

The ICS Caption Editor proves that captioning is not just an accessibility requirement but a pedagogical tool. By involving students in the process, the workload is socialized, and the final output becomes a searchable "videobook" that functions like a textbook.

Limitations & Future Work

The study was based on a smaller sample size (24 students). Additionally, the incentive structure—how to keep students motivated to participate consistently—remains an open question, though many indicated a willingness to work for academic credit. Future iterations could benefit from Large Language Models (LLMs) to provide "context-aware" auto-corrections before the students even begin their review, further reducing the manual workload.


Source Context: This work was supported by the National Science Foundation (NSF) and integrated into the ICS Videos framework used at the University of Houston.

Find Similar Papers

Try Our Examples

  • Search for recent studies on the accuracy improvements of Whisper or other modern transformer-based ASR models on technical STEM lecture datasets compared to the baseline ASR tools used in 2013.
  • Which paper first established the 'Indexing, Captioning, and Searchable' (ICS) framework, and how has the underlying video segmentation logic evolved since the implementation of the ICS Video Player?
  • Explore research that applies crowdsourced transcription methods to real-time or low-latency captioning for live-streamed educational content in Metaverse or VR learning environments.
Contents
ICS Caption Editor: Leveraging Crowdsourcing to Bridge the ASR Accuracy Gap in STEM Education
1. TL;DR
2. The "Broken" State of Automated Transcription
3. Methodology: The Power of the Crowd
3.1. Key Architectural Features:
4. Experimental Results: Accuracy Meets Efficiency
5. Academic Insight: Why it Works
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations & Future Work