Leveraging Wikipedia: Solving Information Overload in the Real-Time Social Web
13261_Semantic Filtering for Social Data.
The paper proposes a novel information-filtering framework for social networks by leveraging Wikipedia as a dynamic, crowd-sourced knowledge base. It introduces the Hierarchical Interest Graph (HIG) and semantic hashtag tracking to overcome the limitations of short-text processing and real-time vocabulary shifts.
TL;DR
Social networks like Twitter and Facebook generate over 5 billion microblogs daily, creating a "poverty of attention." This paper presents a methodology to filter this deluge by using Wikipedia as a live knowledge base. By building Hierarchical Interest Graphs (HIG) and tracking evolving hashtags, the system provides context to cryptic short-texts and adapts to real-world events in near real-time.
Background: The Social Data Dilemma
In the era of "Arab Spring" and "Hurricane Sandy," social media has become the primary mechanism for information dissemination. However, two technical hurdles prevent effective filtering:
- Lack of Context: A tweet saying "Cubs beat Reds" is meaningless to a system that doesn't know these are baseball teams.
- Dynamic Vocabularies: During the 2014 Indian elections, hashtags shifted from #NaMo to #VoteForRG instantly. Static dictionaries cannot keep up.
While Linked Open Data (LOD) like DBpedia offers structure, it lacks the update velocity of the real world. This study turns to Wikipedia, which is updated by the crowd at a pace comparable to news cycles.
Methodology: From Hierarchical Interest to Evolving Semantics
1. Hierarchical Interest Graphs (HIG)
The authors argue that human interest isn't just about keywords; it’s about categories. If you tweet about the "Chicago Cubs," you are likely interested in "Major League Baseball."
The system extracts taxonomic knowledge from Wikipedia to build a HIG. It uses a Spreading Activation Algorithm to assign scores to related topics. This allows the system to filter tweets that don't even mention the original keyword but are semantically relevant.

2. Tracking Evolving Hashtags
To handle the "Real-Time" challenge, the authors use hashtag co-occurrence. By monitoring which hashtags appear together and validating their relationship through Wikipedia’s hyperlink structure, the system can discover new "filter keywords" automatically.

Experiments & Results: Precision in the Chaos
The researchers tested their approach on high-volatility events like the US presidential election and Hurricane Sandy.
- Contextual Filtering: As shown in the table below, the HIG-based profile successfully filtered relevant tweets about Willie McCovey or Sergio Romo for a "Baseball" fan, even when the specific term "Cubs" was missing. Standard keyword filters failed these cases.
- Tracking Accuracy: Their hashtag detection system achieved a Mean Average Precision (MAP) of 0.92, proving that the system can find the "needle in the haystack" even as the needle changes shape.

Critical Insight: The "Tip of the Iceberg"
While highly effective, the paper acknowledges a crucial limitation: The Bottom-Up Lag. In events like terrorist attacks or civil protests, Twitter often moves faster than Wikipedia's consensus-building editors. For these hyper-recent "black swan" events, the knowledge base may still lag slightly behind the raw feed.
Conclusion
This work demonstrates that the gap between "unstructured social noise" and "structured knowledge" can be bridged by crowd-sourced intelligence. By moving from keyword matching to Hierarchical Semantics, we move closer to an information-filtering system that understands what we care about, not just what we say.
Takeaway for Practitioners: When dealing with short-form text, don't just look at the tokens; look at the graph they belong to. Wikipedia isn't just an encyclopedia; it's a real-time map of human interest.
