Automated Large-Scale Data Acquisition: Scaling Crosswalk Classification via Crowdsourcing

Computers & Graphics

2023-01-01
T. Polasek, Martin Čadík, Y. Keller, Bedrich Benes
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an automated system for large-scale crosswalk classification using Deep Learning (VGG-16) and crowdsourced data from OpenStreetMap and Google Street View. By leveraging programmatic data acquisition and annotation, the authors created a massive dataset of over 650,000 images, achieving a SOTA overall accuracy of 94.12% without manual labeling.

TL;DR

Researchers have developed a way to bypass the "manual labeling bottleneck" for crosswalk detection by using OpenStreetMap (OSM) and Google Street View (GSV). By automatically scraping and labeling over 600,000 images, they trained a deep learning model that rivals human-curated datasets, achieving up to 96.3% accuracy and demonstrating remarkable robustness across different vehicle-mounted cameras.

Problem & Motivation: The Small Data Trap

Traditional crosswalk classification systems suffer from geographic parochialism. Most models are trained on small, local datasets—perhaps a single neighborhood or city. When these models face the "wild"—aging paint in Brazil, strong shadows in different latitudes, or different camera heights—they often fail.

The cost of manual labeling for a global-scale solution is prohibitive. The authors' insight was simple yet powerful: Why manually label when thousands of volunteers at OpenStreetMap have already done the work? By linking OSM coordinates with GSV imagery, we can turn the world into a self-annotating training set.

Methodology: The "Zero-Labor" Pipeline

The system follows a sophisticated multi-step acquisition process to ensure the quality of the "automatic" labels:

1. Geographic Splitting and Retrieval

The system uses the Overpass API to query highway=crossing. Because the API has limits on the bounding box size, the authors implemented a recursive splitting strategy (shown below) to ensure no high-density crosswalk area is missed.

Region Splitting Strategy

2. Heading Calculation and Image Alignment

To simulate a driver's perspective (Cockpit View), the system doesn't just download a random crop. It calculates the Heading (α) between sequential GPS points using the atan2 function. This ensures the "camera" is pointed exactly down the road at the crosswalk.

Heading Logic

3. Model Architecture

The backbone of the system is a VGG-16 architecture, pre-trained on ImageNet and fine-tuned on the GSV data. The team compared two models:

  • GSV-FA (Fully-Automatic): Labels strictly from the OSM/GSV pipeline.
  • GSV-PA (Partially-Annotated): A subset refined by a human expert to correct OSM errors.

Experiments & Results: Robustness Over Precision

The results were surprising. While the human-refined model (GSV-PA) was statistically better, the fully automatic model (GSV-FA) was incredibly close.

  • GSV-FA Accuracy: 94.12%
  • GSV-PA Accuracy: 96.30%

Cross-Database Generalization (The Real Test)

To prove this isn't just "overfitting to Google's cameras," the authors tested the models on two entirely different sources: IARA (an experimental autonomous car with a Bumblebee XB3 camera) and GoPro (standard action cam).

IARA & GoPro Cross-Database Performance

The models maintained ~90% accuracy on IARA and GoPro datasets, even though they were trained on GSV imagery. This demonstrates that the diversity of a large, scraped dataset outweighs the noise of automatic labels.

Critical Analysis & Conclusion

Takeaway

The paper effectively demonstrates that the quantity and diversity of data generated by crowdsourcing can compensate for the lower quality of automated labels. This "free" data pipeline allows for the rapid deployment of safety features in countries where road maintenance is inconsistent and proprietary mapping data is scarce.

Limitations

  • Daylight Bias: Current datasets are restricted to Google Street View images, which are predominantly taken during the day.
  • OSM Noise: Inaccuracies in OpenStreetMap (misplaced coordinates) can introduce label noise, though the CNN seems surprisingly robust to this "jitter."
  • Subjectivity of "Far": The model remains sensitive to how far away a crosswalk is before it is "classified" (the temporal boundary problem in video sequences).

Future Outlook

This work paves the way for "on-the-fly" extrinsic calibration and synthetic data augmentation (adding lane markings and signs to images) to further improve robustness for autonomous vehicles.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize OpenStreetMap and computer vision for automated road infrastructure inventory and assessment.
  • What are the latest techniques in cross-domain adaptation for autonomous vehicle perception systems when training on web-scraped data versus real-world sensor data?
  • Investigate how the "Zebra Crossing Spotter" methodology has been extended to real-time mobile and wearable applications for the visually impaired since 2017.
Contents
Automated Large-Scale Data Acquisition: Scaling Crosswalk Classification via Crowdsourcing
1. TL;DR
2. Problem & Motivation: The Small Data Trap
3. Methodology: The "Zero-Labor" Pipeline
3.1. 1. Geographic Splitting and Retrieval
3.2. 2. Heading Calculation and Image Alignment
3.3. 3. Model Architecture
4. Experiments & Results: Robustness Over Precision
4.1. Cross-Database Generalization (The Real Test)
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook