Beyond the Ivory Tower: Crowdsourcing the Discovery of Global CDNs

Increasing the Coverage of Vantage Points in Distributed Active Network Measurements by Crowdsourcing

2014-01-01
Valentin Burger, Matthias Hirth, Christian Schwartz, Tobias Hoßfeld, Phuoc Tran-Gia
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a crowdsourcing-based approach to distributed active network measurements, specifically for analyzing global Content Delivery Networks (CDNs) like YouTube. By recruiting workers from the Microworkers platform to run Java-based measurement scripts, the authors achieved significantly higher geographic and network diversity compared to traditional research testbeds like PlanetLab.

TL;DR

Researchers from the University of Würzburg have demonstrated that traditional academic measurement platforms like PlanetLab are blind to large sections of the modern Internet. By using crowdsourcing to turn regular users into measurement probes, they doubled the number of identified Autonomous Systems (ASes) hosting YouTube servers and proved that CDNs are far more localized at the edge than academic networks suggest.

The Blind Spot of Academic Networks

For years, the networking community relied on PlanetLab—a global network of servers located at universities. While valuable, PlanetLab exists within National Research and Education Networks (NRENs).

The problem? NRENs have world-class peering and high-bandwidth backbones that look nothing like the messy, often congested paths of a residential ISP in Bangladesh or Romania. If you want to understand how a global CDN like YouTube balances its load for the average consumer, measuring from a university campus gives you a skewed, overly optimistic perspective.

Methodology: Putting the 'Crowd' in Measurements

The authors compared two distinct measurement approaches:

  1. The Baseline: 220 nodes in PlanetLab, mostly concentrated in the US and Western Europe.
  2. The Innovation: 247 workers from the Microworkers platform across 32 countries.

Each probe (whether a PhD's server or a worker's laptop) performed a simple but effective task: resolving the actual IP addresses of video servers for specific YouTube URLs. This allowed the researchers to map where Google was actually "hiding" its content—whether in its own massive data centers or tucked away in local ISP caches.

Platform Comparison Table

Table 1: The logistics of Crowdsourcing vs. Social Networks showing the speed and reach of paid micro-tasks.

Insights: The Reality of Edge Caching

The results revealed a massive disparity in how the Internet "looks" to different users:

1. Geographic Discovery

PlanetLab's view was heavily US-centric. It saw 80% of YouTube requests being served by US-based hardware. In contrast, the crowdsourced probes—many of which were in Asia-Pacific and Eastern Europe—found that only 44% of requests went to the US. The rest were served by regional or local caches that PlanetLab simply couldn't see.

2. Autonomous System (AS) Coverage

Perhaps the most striking finding was the AS diversity.

  • PlanetLab identified YouTube servers in roughly 30 ASes, mostly Google-owned (AS15169).
  • Crowdsourcing identified servers in over 60 ASes.

This disparity highlights that Google/YouTube heavily utilizes "Google Global Cache" (GGC) nodes placed directly inside local ISPs. Since ISP users are "closer" to these nodes than academic networks are, only the crowdsourced probes could effectively trigger and identify these edge locations.

AS Distribution Comparison Figure: The long tail of AS coverage discovered via Crowdsourcing vs. the limited scope of PlanetLab.

Critical Analysis: Why This Matters

This paper is a wakeup call for network researchers. It argues that locality isn't just about geographic distance; it's about network topology.

Why Crowdsourcing works for researchers:

  • Speed: Campaigns can be completed in hours rather than weeks.
  • Granularity: You can "target" specific countries or ISPs to fill holes in your data.
  • Real-world Bias: It captures the specific peering and transit bottlenecks that define the average user's experience.

Limitations: The authors acknowledge that using human workers introduces "noise." Workers might use VPNs, have malware, or simply fail to follow instructions. Filtering for "reliable users" (as shown in Table 1) is a necessary step that reduces the raw sample size.

Conclusion: The Future of Distributed Measurement

As CDNs become more distributed and integrated into ISP infrastructures through technologies like Edge Computing, academic-only testbeds will become increasingly obsolete for measuring end-user experience. Crowdsourcing offers a viable, low-cost path to achieving a truly global vantage point, turning the "bottleneck" of the residential ISP into a primary source of architectural insight.

Takeaway: If you are measuring the modern web, your probes must live where your users live.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize crowdsourcing for latency and Quality of Experience (QoE) measurements in 5G or mobile edge computing.
  • What are the primary security and data integrity challenges identified in using anonymous crowdsourced workers for active network measurements?
  • Find papers comparing the current Google/YouTube CDN infrastructure (post-2020) with the 3-tier hierarchy identified in early 2010s research.
Contents
Beyond the Ivory Tower: Crowdsourcing the Discovery of Global CDNs
1. TL;DR
2. The Blind Spot of Academic Networks
3. Methodology: Putting the 'Crowd' in Measurements
4. Insights: The Reality of Edge Caching
4.1. 1. Geographic Discovery
4.2. 2. Autonomous System (AS) Coverage
5. Critical Analysis: Why This Matters
6. Conclusion: The Future of Distributed Measurement