Scaling Sentiment: Parallel Label Propagation on Spark GraphX

Large Scale and Parallel Sentiment Analysis Based on Label Propagation in Twitter Data

2018-08-01
Yibing Yang, M. Omair Shafiq
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a scalable and parallel sentiment analysis framework for Twitter data using the Label Propagation Algorithm (LPA). By leveraging the Spark GraphX API, the authors implement a semi-supervised approach that integrates lexicon seeds and social network structures to achieve significant performance gains over traditional baselines.

TL;DR

This paper tackles the challenge of sentiment analysis in the era of Big Data. By moving away from rigid supervised learning and toward Label Propagation on a graph, the authors utilize the Spark GraphX framework to process millions of tweets in parallel. Their approach effectively turns a small set of emoticons and lexicon words into "seeds" that infect an entire social graph with sentiment labels, outperforming traditional lexicon baselines by over 7%.

Background: The Twitter Bottleneck

Sentiment analysis on Twitter is notoriously difficult due to three factors:

  1. Volume: Hundreds of millions of tweets daily.
  2. Style: Slang, typos, and "Twitter-speak" break traditional NLP parsers.
  3. Data Scarcity: Lack of manually labeled data for supervised classifiers.

While MapReduce serves general data tasks, it struggles with the iterative nature of graph algorithms. The authors position their work at the intersection of Parallel Graph Computing and Semi-Supervised Learning, using the graph's structure to bypass the need for massive human-labeled datasets.

The Motivation: From Text to Graph

Why use a graph? Because sentiment isn't isolated. If a user retweets a positive sentiment or uses the same hashtag as a known positive tweet, there is a high probability they share that sentiment. The authors' insight is to represent these relationships (User-Tweet, Tweet-Hashtag, Tweet-Word) as edges in a massive heterogeneous graph.

Methodology: The Label Propagation Pipeline

The core of the system is the Label Propagation Algorithm (LPA). In this model, labels are like a fluid that flows through the edges:

  1. Seed Injection: A small number of nodes are pre-labeled using a sentiment lexicon (e.g., OpinionFinder) or "noisy" labels like emoticons (e.g., :D is positive).
  2. Iterative Flow: In each "superstep," every node updates its label based on the labels of its neighbors.
  3. Convergence: This continues until the labels across the graph stabilize.

Graph Architecture

The graph is built using several layers of features:

  • User Nodes: Connected via "following" and "retweet" relationships.
  • Tweet Nodes: The primary entities being classified.
  • Feature Nodes: N-grams, Hashtags, and Emoticons that link different tweets together.

Overall Sentiment Graph Structure

Parallel Execution with GraphX

To handle millions of nodes, the authors implement this on Spark GraphX using the Pregel API. By separating the graph into partitions, the updates are computed simultaneously across a cluster.

Experiments & Results

The authors tested their system on the Sentiment 140 (large-scale performance) and HCR (accuracy) datasets.

1. Scaling Performance

The results confirm that the system scales well. As shown in the performance chart, increasing the number of worker nodes leads to a nearly linear reduction in processing time initially, before hitting a communication overhead bottleneck at 6-7 nodes.

Speedup of different nodes

2. Accuracy Comparison

The LPA method (61.9%) significantly beat the Lexicon-based baseline (54.2%). Interestingly, the authors found that adding complex social networking edges (like user following) actually decreased accuracy slightly (to 60.6%) compared to using purely textual feature edges. This suggests that "who you follow" is a noisier indicator of sentiment than "what words you use."

Accuracy Comparison Table

Critical Analysis & Conclusion

Takeaway

Label Propagation is a powerful "force multiplier" for sentiment analysis. By using just a few emoticons as seeds, the algorithm can label millions of tweets without human intervention. The use of GraphX makes this feasible for production-level big data.

Limitations

  1. Domain Sensitivity: The model performed worse on the HCR (Healthcare Reform) dataset because the language was more "serious" and sarcastic, illustrating that LPA still relies on the quality of the initial seeds.
  2. Social Noise: The finding that social network links (user following) didn't help suggests that sentiment is highly context-specific and doesn't always follow social clusters.

Future Work

The authors propose adding Community Detection before classification. By segmenting the graph into topical clusters first, the label propagation could be restricted to relevant sub-graphs, potentially filtering out noise and improving accuracy in specialized domains.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) instead of Label Propagation for large-scale Twitter sentiment analysis on Spark or Flink.
  • Which paper first proposed the "Modified Adsorption" (MAD) algorithm for graph-based learning, and how does it compare to the original Label Propagation implementation used in this study?
  • Investigate how community detection algorithms like Louvain or Infomap have been combined with sentiment analysis to improve classification accuracy in domain-specific social media datasets.
Contents
Scaling Sentiment: Parallel Label Propagation on Spark GraphX
1. TL;DR
2. Background: The Twitter Bottleneck
3. The Motivation: From Text to Graph
4. Methodology: The Label Propagation Pipeline
4.1. Graph Architecture
4.2. Parallel Execution with GraphX
5. Experiments & Results
5.1. 1. Scaling Performance
5.2. 2. Accuracy Comparison
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Work