CRR: Breaking the "Independence" Barrier in Video Random Access
Crowdsourcing Based Cross Random Access Point Referencing for Video Coding
This paper introduces Cross Random-access-point Referencing (CRR), a novel video coding structure that allows Random Access Point (RAP) pictures to utilize External Reference Pictures (ERPs) from non-adjacent segments. By integrating a crowdsourcing-based optimization for ERP selection, CRR achieves state-of-the-art gains of 12.00% on VVC common test sequences and up to 25.48% on long-duration drama content.
TL;DR
To support jumping to different timestamps (Random Access), modern video encoders like VVC insert "Intra" frames that don't look back at previous data. This wastes bits because the same background often reappears later. Cross Random-access-point Referencing (CRR) solves this by allowing these entry points to look "sideways" at shared External Reference Pictures (ERPs). Using Crowdsourcing Theory, the authors optimize which pictures to store as "references" for the whole video, slashing bitrates by up to 25% for long-form drama series.
Problem & Motivation: The "Amnesia" of Modern Encoders
In current standards (HEVC, VVC), a video is divided into Random Access Segments (RASs). To ensure you can start playing from the middle of a file, each RAS begins with an IDR or Intra-RAP picture.
- The Pain Point: These segments are independent. If a character is talking in Scene A, then we cut to Scene B, then back to Scene A, the encoder "forgets" Scene A and must re-encode it from scratch at the next Random Access Point.
- The Gap in Prior Work: Methods like DRAP (Dependent RAP) allowed some dependency but only on the immediately preceding intra frame. This is useless for episodic content or news where Scene A and Scene C are similar, but Scene B is totally different.
Methodology: Crowdsourcing the Best References
The authors treat the selection of reference frames as an optimization problem. Instead of just picking the first frame of a segment, they look for External Reference Pictures (ERPs) that provide the best "deal" for the entire video.
1. The CRR Structure
Unlike conventional structures where the Decoded Picture Buffer (DPB) is cleared at every RAP, CRR maintains an "External DPB."
In Fig 1(c), you can see multiple ERPs being used across non-consecutive segments, breaking the linear chain of traditional coding.
2. Crowdsourcing ERP Selection
How do you pick which frames should be ERPs? If you pick too many, the "cost" (bitrate of the ERPs themselves) outweighs the "contribution" (savings in the RASs). The authors map this to Crowdsourcing Theory:
- Tasks: Compressing the video segments.
- Users: Candidate ERP frames.
- Profit Function: The reduction in RD-cost minus the cost to transmit the ERP.
By proving the function is submodular (the law of diminishing returns applies), they use a Local-Search-Based (LSB) algorithm to iteratively add or remove ERPs until the total bitrate is minimized.
Experiments & Results: Massive Wins for Long-Form Content
The authors tested CRR against VVC (VTM 3.0) and various DRAP flavors.
SOTA Comparison
On standard "Common Test Condition" (CTC) sequences, CRR gained 12.00%. However, the real power shows in 2-minute drama clips (like Sherlock or The Big Bang Theory):
- Average Gain: 25.48% BD-rate saving.
- Efficiency: In the movie "The Man from Earth," CRR effectively identified scene clusters, allowing the encoder to reuse background data across 37 different fragments using only a handful of ERPs.

System Integration
One of the paper's highlights is the practical design for DASH (Streaming over HTTP). They propose a separate "External Track" for ERPs. When a user seeks to a random point, the player fetches the necessary ERP just seconds before the main segment, ensuring "Random Access" functionality remains intact without needing a massive local buffer.
Critical Analysis & Conclusion
Takeaway
CRR transitions video coding from a "forgetful" stream to a "context-aware" library. By treating ERP selection as a crowdsourcing problem, it provides a mathematically sound way to balance storage overhead vs. compression gain.
Limitations
- Encoding Complexity: CRR adds about 12% to the encoding time. While small for offline VOD, it might be too heavy for live broadcast.
- Content Sensitivity: If a video has zero repeating scenes (e.g., constant fast motion in a forest), CRR might actually introduce a slight overhead.
Future Prospect
This work lays the groundwork for "Cloud-based Reference Libraries," where a streaming service could maintain a shared set of background reference frames for an entire series (e.g., the "Central Perk" set in Friends), used by all episodes to achieve unprecedented compression levels across the whole library.
