Cross-Platform Triangulation: Leveraging Big Data to Detect Social Media Rumors
Detecting False Information of Social Network in Big Data
This paper introduces a novel cross-platform validation model for social network false information detection. It converts social network posts and search engine results into three-dimensional vectors (event, time, place) and calculates similarity and emotional consistency to determine authenticity, achieving an 88.33% precision rate.
TL;DR
Researchers from the Beijing University of Posts and Telecommunications have developed a model that detects false information by comparing social media posts against authoritative web data. Instead of just looking at who posted the news, the model analyzes what happened by converting events into 3D vectors (Event, Time, Place) and verifying them via Google-screened authoritative websites.
Key Achievement: Reached an 88.33% Precision rate, significantly outperforming methods that rely solely on user reputation or internal sentiment.
The Problem: The "Reputation Trap" in Rumor Detection
In the era of Big Data, misinformation spreads faster than truth. Most existing detection systems fall into two traps:
- The User-Centric Bias: They assume posts from "high-credibility" users are true. However, even influential accounts can be hacked or misled.
- Internal Sentiment Analysis: They look for "panic" within the post itself but ignore whether the event actually occurred in the real world.
The authors argue that the only way to guarantee authenticity is through external verification—matching social claims against the broader "Internet consensus."
Methodology: The 3D Event Vector & Consistency Layer
The proposed model follows a sophisticated pipeline to move from a single post to a truth verdict.
1. 3D Vector Conversion
Every piece of information is distilled into a vector :
- (Event Name): What happened?
- (Time): When did it happen?
- (Place): Where did it occur?
2. Information Screening (WQ Value)
The model queries Google for the social media keywords but doesn't trust all results equally. It calculates a Website Quality (WQ) Value: This ensures that information from established portals (like official news sites) carries more weight than personal blogs.

3. Semantic Similarity and Consistency
Using HowNet Semantic Information, the model calculates the similarity between the social post's vector and the Internet's vectors. Crucially, it adds a Consistency Detection step. If a social post says "A bomb exploded" (negative sentiment) and an Internet source says "The bomb report was a drill" (informational/neutral), the emotional orientation mismatch flags the social post as false, even if the keywords "bomb" and "place" match.
Experimental Results
The model was tested against the Sina Microblog hot events dataset.
| Method | Precision | Recall | F-Measure |
|---|---|---|---|
| Proposed Model | 88.33% | 93.33% | 90.76% |
| User Credibility Mode | 69.01% | 70.04% | 69.52% |
| Emotion/Opinion Mode | 86.10% | 86.00% | 86.05% |
The results clearly show that comparing content against the "Internet truth" is far more effective than analyzing user background. The Recall of 93.33% is particularly impressive, suggesting the model is excellent at catching rumors that others miss.
Figure: The impact of adding the consistency detection module—notice the sharp rise in precision when sentiment orientation is considered.
Critical Insight & Future Outlook
The genius of this work lies in its Inductive Bias: the belief that collective Internet data (verified by PageRank) serves as a "ground truth" for ephemeral social media claims.
Limitations: The current model relies heavily on Google search results. As the authors admit, if an event is so new that it hasn't been indexed, or if search results are manipulated, the model's performance may degrade.
Conclusion: This research moves us away from studying "who is talking" and toward "what is being said." By quantifying events into mathematical vectors and checking them against a weighted web of trust, we can finally build a more objective shield against the spread of false information.
