Intelligent Data Augmentation: Streamlining VQA Dataset Expansion via Crowdsourcing
A Crowdsourcing Tool for Data Augmentation in Visual estion Answering Tasks
This paper introduces a specialized crowdsourcing tool designed for data augmentation in Visual Question Answering (VQA) tasks, specifically focusing on binary (yes/no) questions. By integrating Siamese Networks, SSD object detection, and Word2Vec, the tool intelligently filters candidate images and questions to optimize the human curation process.
TL;DR
The paper presents a collaborative crowdsourcing framework designed to tackle the "data hunger" of Visual Question Answering (VQA) models. By utilizing deep learning modules (Siamese Networks, SSD, and Word2Vec) to pre-filter and rank image-question pairs, the tool reduces the curation workload by over 1,000 times compared to brute-force labeling, specifically for binary VQA tasks.
Background and Motivation
Visual Question Answering is the "holy grail" of multi-modal AI, requiring a model to understand both visual scenes and natural language nuances. However, the state-of-the-art—Deep Neural Networks—suffers from a harsh reality: gaining a 10-12% boost in accuracy often requires a 10x increase in dataset size.
Manual labeling at this scale is prohibitively expensive. While "blind" data augmentation (like flipping images) exists, it often leads to redundant data that offers diminishing returns. The authors identified a crucial need for a "disciplined augmentation" approach—one that finds new, relevant images from external sources (like ImageNet) and matches them with appropriate questions to be verified by humans.
Methodology: The Four-Stage Pipeline
The system's architecture is built on a series of intelligent filters designed to maximize the "value per click" for human curators.
1. Image Filtering (The Siamese Gate)
To ensure the new images () are relevant to the original task (), the tool uses a Siamese Neural Network. This network computes a similarity score () between pairs of images. Only images that cross a similarity threshold () move forward.
2. Semantic Extraction
The tool then performs a dual-modality analysis:
- Visual Side: Uses Single Shot Detection (SSD) to identify objects within the new images.
- Textual Side: Uses a POS Tagger to extract nouns from existing questions in the original dataset.
3. Question Filtering (Word Embeddings)
Using Word2Vec, the tool measures the semantic distance () between the objects identified in the image and the nouns in the question. If a question is semantically "too far" from the image content, it is discarded.
4. Prioritized Curation
The remaining pairs are presented to humans, ranked by their relevance. The curator simply confirms if a question applies to the image and provides the final binary answer.
Figure 1: The architecture of the curation tool, showcasing the modular flow from raw data to human verification.
Experimental Results
In a practical instantiation targeting "Dog" and "Cat" categories from ImageNet and the VQA dataset, the numbers were striking:
- Initial Candidate Pairs: ~341 Million
- Post-Filtering Pairs: 259,716
- Efficiency Gain: ~1,300x reduction in human effort.
- Human Performance: Curators took about 12 seconds per item, suggesting a highly focused and low-friction interface.
Figure 2: An example of the SSD model detecting objects (Bike, Car, Dog) to generate terms for semantic matching.
Critical Insight & Future Outlook
The brilliance of this work lies in its Search-Space Reduction. Instead of asking humans to "write a question for this image" (high cognitive load), it asks them to "verify if this existing question fits this similar image" (low cognitive load).
Limitations & Future Work
- Binary Focus: The current tool is limited to "Yes/No" questions. Expanding to open-ended questions remains a challenge due to the increased complexity of answer verification.
- Active Learning: The authors plan to incorporate Active Learning to prioritize not just "similar" images, but those that the VQA model is currently "uncertain" about.
Conclusion
This tool represents a significant step toward scalable VQA dataset creation. By treating data augmentation as an intelligent filtering problem rather than a generative one, it paves the way for high-quality, large-scale multimodal datasets that are both semantically diverse and labor-efficient.
