Stance Mining in the Wild: Navigating Skewed Data with Minimalism

Topic-Based Stance Mining for Social Media Texts

2015-01-01
Wei-Fan Chen, Yann-Hui Lee, Lun-Wei Ku
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a two-step approach for Topic-Based Stance Mining on social media (Facebook), specifically targeting the controversial "Anti-reconstruction of Nuclear Power Plant" issue. It utilizes InterestFinder for relevance filtering and integrates the compositional sentiment tool CopeOpi with a polarity-to-stance transition mechanism to achieve high-precision classification.

TL;DR

Analyzing public opinion on Facebook is notoriously difficult due to the "needle in a haystack" problem—valuable minority opinions are often buried under mountains of neutral or skewed data. This paper introduces a two-step mining approach that starts with just five seed words. By leveraging a sentiment analysis tool (CopeOpi) and a logical transition process, the researchers achieved up to 94.62% precision in stance classification, outperforming traditional SVM and co-training methods in low-label environments.

Problem & Motivation: The Skewness Trap

Most sentiment analysis research assumes a relatively balanced dataset (e.g., 50/50 split of positive/negative movie reviews). However, real-world political discourse is messy:

  1. Extreme Imbalance: In the "Anti-reconstruction of Lungmen Nuclear Power Plant" case study, "unsupportive" evidence made up only 1.25% of the data.
  2. The Polarity/Stance Gap: A user might write a "negative" post criticizing the "unsupportive" party. A naive sentiment tool would label this "negative," but the actual stance is "supportive."
  3. Label Scarcity: New social issues emerge overnight. We cannot wait for thousands of manually labeled examples before we start analyzing them.

Methodology: From Seed Words to Stance

The authors propose a logic-driven pipeline rather than a "black box" machine learning approach.

1. Relevance Filtering (InterestFinder)

Before classifying stance, the system must filter out noise. Using InterestFinder, the system identifies key terms related to five seed words (e.g., "abandon nuclear," "nuclear power"). Only documents with a high topical interest score are passed to the next stage.

2. The Co-training Attempt

The researchers initially experimented with an SVM co-training process. The idea was to have two "viewpoints" train each other:

  • Content View: Bag of Words (BOW), Word Vectors (GloVe), and Dependency Trees.
  • Metadata View: Social signals like Likes, Shares, and Comments.

SVM Co-training Comparison Table 4: Showing results of SVM co-training with BOW features.

3. The Core Intuition: Polarity Transition

The breakthrough came from using CopeOpi, a compositional sentiment tool. Instead of just looking at the post polarity, they categorized the seeds themselves:

  • SUP Seeds: (e.g., "Embracing nuclear")
  • UN_SUP Seeds: (e.g., "Anti-nuclear")

By calculating the sentiment towards these specific aspects, they derived a Standpoint Score (STD_PT): STD_PT = SUP_Sentiment - UN_SUP_Sentiment + Neutral_Sentiment

This "transition" allows the system to understand that having a positive sentiment toward an "Anti-nuclear" word actually indicates an unsupportive stance toward the power plant.

Experiments & Results

The researchers tested their method on 41,902 Facebook posts. While pure SVM failed to capture the minority class effectively due to the skewness, the CopeOpi + Transition Process reached remarkable precision levels.

Final Performance Results Table 12: Final performance showing the superiority of the transition-based approach.

Key Findings:

  • Precision is King: The high precision (94.62% for Supportive) ensures that decision-makers receive accurate evidence.
  • Dependency over BOW: When using machine learning, combining Word Vectors with Dependency Tree relations (to focus on "root" verbs/actions) was more effective than simple Bag of Words.
  • Rule-based Wins in Skewed Data: In scenarios with only 51 labeled "unsupportive" examples, the logical mapping of sentiment-to-aspect outperformed statistical learners.

Critical Insight & Conclusion

This paper serves as a vital reminder that logical domain knowledge (aspect mapping) often trumps model complexity when data is sparse and imbalanced. While recent LLMs might handle these tasks via prompt engineering today, the fundamental challenge of mapping sentiment to standpoint remains a core hurdle in computational linguistics.

The primary limitation identified is the difficulty in identifying the extreme minority (the 1.25%). Future work should focus on active learning or synthetic data generation (like SMOTE or LLM-based augmentation) to better bridge the gap for these critical minority voices.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Zero-shot or Seed-word based stance detection in social media contexts similar to political or social movements.
  • Which study first introduced the CopeOpi sentiment analysis system, and how has it been modified for modern transformer-based architectures?
  • Find research that investigates the use of multi-modal features (metadata and text) for imbalanced classification in stance mining tasks on platforms like Facebook or X (Twitter).
Contents
Stance Mining in the Wild: Navigating Skewed Data with Minimalism
1. TL;DR
2. Problem & Motivation: The Skewness Trap
3. Methodology: From Seed Words to Stance
3.1. 1. Relevance Filtering (InterestFinder)
3.2. 2. The Co-training Attempt
3.3. 3. The Core Intuition: Polarity Transition
4. Experiments & Results
5. Critical Insight & Conclusion