Beyond the Follow Button: Classifying User Intent via Issue Clusters and Influential Supporters
Follower Classification Based on User Behavior for Issue Clusters
The paper introduces a novel follower classification framework that categorizes Twitter users into supporters, non-supporters, or neutrals relative to a target "authority" user. The core approach utilizes an "Issue Cluster" methodology combined with a modified HITS algorithm to identify influential supporters and track sentiment alignment over specific "trust periods."
TL;DR
On platforms like Twitter, a "follow" does not always equate to "support." This paper proposes a robust classification system that identifies whether followers are supporters, non-supporters, or neutral by analyzing Issue Clusters and Influential Supporters. By moving beyond simple text-based SVMs and looking at the network dynamics of retweets during specific "trust periods," the authors achieved a 20.3% increase in classification accuracy.
Background: The "Follower" Ambiguity
In the landscape of Social Network Services (SNS), authority users—such as politicians or celebrities—amass followers with vastly different agendas. Some follow to cheer, others to criticize, and many just to observe. Standard sentiment analysis often fails here because tweets are short and frequently lack explicit emotional keywords. The authors argue that to truly understand a follower, we must look at who they retweet and when they do it in relation to specific social issues.
Methodology: The Three Pillars of Classification
1. Identifying the Inner Circle (Influential Supporters)
The authors utilize a modified HITS (Hyperlink-Induced Topic Search) algorithm. They define "Influential Supporters" as users who consistently amplify the target user's message.
- Logic: If User A retweets the target user frequently, and User B retweets User A, User A gains "Authority" and "Hub" scores.
- Expansion: These influential supporters act as an "expanded concept" of the target user. If a random follower retweets an influential supporter, they are likely a supporter of the original target user.
2. The Issue Cluster & Trust Period
Public opinion is volatile. An issue relevant today might be forgotten in two weeks. The authors introduce the Trust Period:
- Start: The first time a target user mentions a keyword.
- End: The last time that keyword is retweeted by an influential supporter.
By clustering tweets within this window using a specialized tf-idf weight (multiplied by a "supporter frequency" score), the system identifies the "pulse" of an issue.

3. Resolving the "Co-Follower" Conflict
Many users follow multiple rival politicians. To solve this, the authors use a Bias Ratio: This mathematical approach quantifies relative loyalty, acknowledging that support is often comparative rather than absolute.
Experimental Results
The study analyzed 10,000 followers for five prominent Korean politicians (e.g., Moon J.I., Park G.H.).
| Method | Accuracy | Precision | F1 Measure |
|---|---|---|---|
| SVM (Baseline) | 0.565 | 0.640 | 0.590 |
| Partial (HITS only) | 0.631 | 0.681 | 0.648 |
| Proposed (Issue Clusters) | 0.680 | 0.705 | 0.675 |
The results clearly indicate that incorporating Issue Clusters provides a significant boost over purely linguistic models (SVM). The network density (shown in Figure 3) illustrates how the Issue Cluster expansion captures a much wider array of classifiable interactions than traditional methods.
Fig 3. Visualization of the expanded user network using the proposed methodology.
Critical Insight & Conclusion
The brilliance of this work lies in its Inductive Bias: it assumes that user behavior (retweeting) within a specific temporal context (trust period) is a stronger signal of sentiment than the text itself.
Takeaway: For developers and researchers building social listening tools, the lesson is clear—don't just analyze what is said. Analyze who is amplifying whom during the peak of a specific controversy. While the 68% accuracy leaves room for improvement (potentially through modern LLMs), the framework of using influential cohorts to label data is a powerful strategy for handling sparse text data.
Limitations: The system relies heavily on retweet behavior. "Silent followers" (lurkers) remain difficult to classify. Future work would benefit from incorporating "like" patterns or dwell time if that data were available via API.
