ASTC & ASTCx: Unveiling Social Communities through the Lens of Topics and Sentiments
Community discovery using social links and author-based sentiment topics
The paper introduces ASTC and ASTCx, two probabilistic generative models for community discovery that integrate social links, author-recipient interactions, and sentiment-topic distributions. ASTCx specifically decouples sentiment and topic vocabularies, achieving superior interpretability in uncovering "Topic-Sentiment Unambiguous Communities" across the Enron and Twitter datasets.
TL;DR
In the era of social networking, a community isn't just a set of links; it's a shared conversation with distinct emotional undertones. This paper presents ASTC and ASTCx, two generative models that go beyond network topology to discover communities by blending social links with author-specific sentiment topics. By separating "what" people talk about from "how" they feel, the authors provide a clearer map of social group dynamics.
Background: Why Links Aren't Enough
For years, community detection was a graph theory problem—finding dense clusters of nodes. However, in a corporate email network like Enron or a volatile space like Twitter, a link only tells half the story. Two people might communicate frequently but hold diametrically opposite views on a topic, potentially belonging to different "sentiment communities." Prior works integrated content (topics), but they often ignored the sentiment bias that defines social cohesion or friction.
Methodology: The ASTCx Architecture
The authors propose two versions of their model. While ASTC mixes sentiment and topic words in a single distribution, ASTCx (the extended version) introduces a vital refinement: it treats sentiment words and topic words as separate entities.
1. The Generative Intuition
The model assumes that for every document:
- A Community is sampled.
- An Author and their Recipients are sampled based on that community.
- For every word, a Topic is chosen based on the author's preference within that community.
- A Sentiment label is then sampled, conditioned on that specific topic.
2. De-coupling Content and Emotion
By using a subjectivity lexicon (like MPQA) and WordNet, ASTCx separates adjectives and adverbs. This allows the model to learn that "iphone" and "nexus" are topic words, while "amazing" and "terrible" are sentiment markers, preventing the "topic noise" from diluting the "sentiment signal."
Figure 1: Plate notation for the ASTC model, showing the dependency between communities, authors, topics, and sentiments.
Experimental Insights
The researchers validated their models on the Enron email dataset and the Sanders-Twitter Sentiment Corpus.
Diverse Roles in Single Communities
A fascinating finding was the analysis of "active authors." In the Enron dataset, a single community might discuss multiple topics, but different authors within that community have different "dominant" topics. For example, a Vice President (Steven Kean) showed a broad distribution across all topics, reflecting his managerial oversight, while technical heads were more specialized.
The Clarity of ASTCx
The superiority of the ASTCx model is most visible in its output tables. Traditional models produce a "soup" of words. ASTCx produces structured "Sentiment-Topic" clusters:
- Topic (Google): Positive words: "Sweet", "Nice"; Negative: "Sold", "Wrong".
- Topic (California Energy): Positive words: "Successful", "Strategic"; Negative: "Crisis", "Refunds".
Table 1: Mixed topic-sentiment words from the basic ASTC model.
Table 2: Refined sentiment words for selected topics in the ASTCx model, showing much higher readability.
Professional Opinion: Bridging the Gap
The core value of this work lies in its Inductive Bias. By assuming that sentiment is topic-dependent (e.g., your sentiment towards "Microsoft" may differ from your sentiment towards "Apple"), the model mirrors human psychology more accurately than generic clustering.
Limitations
- Data Sparsity: As the authors noted, the author-recipient relationship is often sparse. In a massive network, many users might only interact once, making the Dirichlet priors difficult to settle.
- Lexicon Dependency: ASTCx relies on external tools like WordNet. Its performance is capped by the quality of these linguistic resources.
Conclusion
ASTCx provides a robust framework for managers and social analysts to not only see who is talking to whom, but to understand the core ideological or emotional consensus of those groups. Future directions, such as automatically determining the number of communities (M) and topics (K), will be essential for scaling this to the "Big Data" of modern social media.
