Crowdsourcing Highlights: Extracting Sports Events via Social Media Bursts
A Framework for Extracting Sports Video Highlights Using Social Media
The paper proposes a novel framework for sports video highlight extraction by leveraging social media text streams (e.g., PTT posts) instead of traditional video analysis. By treating comment volume as a time series, the system detects events via "spikes" and annotates them using a weighted TF-IDF approach, achieving results comparable to human-generated summaries for the 2014 FIFA World Cup.
TL;DR
Instead of using computationally expensive Computer Vision to "watch" sports videos, this paper proposes "listening" to the crowd. By analyzing spikes in social media posts (like the PTT forum during the FIFA World Cup), the researchers can pinpoint and label major game highlights. The result? A highly efficient, scalable summarization tool that mirrors human expertise by focusing on "Power Users."
Problem & Motivation: The Computational Burden of Vision
Extracting highlights from a 90-minute football match is a classic problem in multimedia. Traditionally, researchers have taken two paths:
- Manual Labeling: Accurate but impossible to scale in the age of "explosive" video data.
- Content-Based Analysis: Using AI to detect a ball hitting a net or a referee’s whistle.
While the second approach is automated, it suffers from the "Semantic Gap"—it is hard for a computer to understand why a certain play is emotionally significant or "highlight-worthy" just from pixels. Moreover, processing high-resolution video frame-by-frame is computationally draining.
The authors' insight is simple: Social media users are real-time sensors. When something important happens, the volume of posts spikes. These "bursts" are the keys to the game's pulse.
Methodology: From Spikes to Semantics
The proposed framework operates in three distinct stages:
1. Event Detection: Finding the "Spikes"
The authors analyzed the time-series data of comment volumes. They proposed two methods to find the "magic moments":
- Total Amount (TA): Simple ranking by post volume per minute.
- Moving Average (MA): Detecting a sudden surge relative to the recent past, which is often more sensitive to unexpected events like a sudden red card.

2. Semantic Annotation: The "Power User" Filter
Identifying when an event happened is only half the battle; we also need to know what happened (e.g., a "Goal" vs. a "Yellow Card"). While TF-IDF is a standard tool for finding important keywords, it is often cluttered with "noise"—casual chatter or unrelated slang.
The authors introduced a User Weighting mechanism. They identified "Power Users"—those whose posting patterns align closely with the detected event bursts. By performing TF-IDF only on the comments of these high-weight users, the system ignores the "white noise" of the crowd and captures the technical semantics of the game.

Experiments: Testing on the World Cup 2014
The team tested their framework on iconic matches, including Germany vs. Brazil and the 2014 Final.
- MA vs. TA: The Moving Average method proved more robust, showing higher precision and recall across most games.
- The Power of the Few: When isolating the top keywords, the "Power User" method consistently outperformed the "All Users" method. This confirms that a small subset of engaged users provides the most accurate descriptions of live events.

Critical Analysis & Conclusion
Takeaway
This framework shifts the paradigm of video analysis from Computer Vision to Social Data Mining. It treats human observers as a biological "feature extractor," drastically reducing the cost of video summarization while maintaining high semantic accuracy.
Limitations
While effective, the method is dependent on the synchronicity and volume of social media. For niche sports or matches with low social engagement, the "spikes" might be too weak to detect. Furthermore, there is a natural "delay" between a live event and the social media reaction, which requires alignment calibration.
Future Outlook
As Large Language Models (LLMs) become more efficient, replacing basic TF-IDF with an LLM to summarize the "Power User" comments could lead to even more nuanced and descriptive highlight reels, moving from keyword tags to full narrative summaries.
