Automatic vs. Human: Solving the Ambiguity of "Relevance" in Social Media

Human vs. Automatic Annotation Regarding the Task of Relevance Detection in Social Networks

2018-01-01
Nuno Guimarães, Filipe Miranda, Álvaro Figueira
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates methodologies for building large-scale datasets to detect "journalistic relevance" in social media posts across Facebook and Twitter. It proposes and compares a multi-filter Human Annotation approach via crowdsourcing against an Automatic Assessment method that labels posts based on their similarity to real-time professional news RSS feeds.

TL;DR

Is a post about a local fire "relevant" to the world? Is a celebrity tweet "newsworthy"? Human annotators often disagree based on personal interest. This paper demonstrates that automatic labeling based on similarity to professional news sources creates better training data for AI than filtered crowdsourcing, achieving a superior F1-score of 0.64 in detecting journalistic relevance.

Contextual Positioning

Within the landscape of social media analysis, we've moved past simple event detection (like tracking an earthquake in real-time). The current challenge is the broad-spectrum detection of newsworthiness. This work positions itself as a methodology study, proving that for ambiguous concepts like "relevance," an objective "silver standard" (automatic) beats a noisy "gold standard" (human).

The Core Problem: The Subjectivity Trap

Supervised learning requires a Ground Truth. However, in the context of social networks:

  1. Ambiguity: What is relevant to a sports fan isn't relevant to a political analyst.
  2. The Crowd Problem: Low-paid workers often rush tasks, providing noisy data.
  3. Cost: Getting 5+ workers per post to reach consensus on 10,000+ posts is economically unfeasible.

Methodology: Two Paths to Relevance

The researchers compared two distinct ways to label ~10,000 social media entries (Tweets and Facebook posts).

1. The Human Approach (Crowdflower + Intelligent Filtering)

To avoid the cost of multiple annotators, they used one worker per post but applied a sophisticated filtering pipeline:

  • User Agreement Rate Deviation: A Bayesian estimate was used to penalize users who consistently diverged from the majority in previous tasks.
  • Consistency Checks: Using Levenshtein Distance to ensure workers weren't just typing random characters in summary fields.
  • News-Awareness Consistency: Identifying users whose self-reported news knowledge varied wildly during the session.

2. The Automatic Approach (News-Similarity)

This was the "Automated" path. If a post shared entities (people, locations, organizations) and keywords with an actual news article published on the same day by entities like The New York Times or CNN, it was automatically labeled as "Relevant."

Crowdsourcing Quality Verification Figure 1: Analyzing user consistency to filter out low-quality human contributors.

Experiments & Results: Automation Wins

The team extracted a variety of features including Textual (pronouns, verb tenses), Sentiment, and Entity Characteristics (using the Guardian API to see how "controversial" or "frequently mentioned" an entity was).

When tested against a dataset labeled by experts (professional standard), the results were clear:

AlgorithmAutomatic Annotation (F1)Human Annotation (F1)
Naive Bayes0.640.59
Gradient Boosted0.570.55
SVM0.500.28

Experimental Results Figure 2: Performance comparison across different machine learning models.

Why did Automation win?

The authors argue that human workers—even when trusted—cannot separate their personal interest from journalistic value. If a worker doesn't care about the "Champions League," they label it "Irrelevant," even though it is objectively newsworthy. The automatic system, by tethering labels to actual news agency output, bypasses this cognitive bias.

Critical Insights & Future Outlook

  • The Strength of External Knowledge: This paper highlights that for social media, "context" doesn't just come from the text, but from the surrounding world (the news cycle).
  • Limitations: The automatic system is limited by its news sources. If a local event is relevant but hasn't reached major RSS feeds yet, it will be mislabeled as "not relevant" (False Negative).
  • The Next Frontier: The authors suggest integrating automatic fact-checking to ensure that "relevant" posts aren't just "fake news" spreading rapidly.

Takeaway for Data Scientists

When building training sets for subjective tasks, look for surrogate objective signals (like RSS feeds or professional databases) before defaulting to expensive and potentially biased crowdsourcing. Noise in the "silver label" is often easier to handle than the fundamental bias in human "gold labels."

Find Similar Papers

Try Our Examples

  • Find recent studies that utilize "distant supervision" or automatic labeling for multi-domain news relevance detection in social media beyond 2017.
  • Which paper first introduced the Bayesian estimate approach for filtering unreliable crowdsourcing workers, and how has it been evolved for NLP tasks?
  • How do modern Large Language Models (LLMs) compare against entity-matching systems when used as zero-shot evaluators for "journalistic relevance" in short-form text?
Contents
Automatic vs. Human: Solving the Ambiguity of "Relevance" in Social Media
1. TL;DR
2. Contextual Positioning
3. The Core Problem: The Subjectivity Trap
4. Methodology: Two Paths to Relevance
4.1. 1. The Human Approach (Crowdflower + Intelligent Filtering)
4.2. 2. The Automatic Approach (News-Similarity)
5. Experiments & Results: Automation Wins
5.1. Why did Automation win?
6. Critical Insights & Future Outlook
7. Takeaway for Data Scientists