The "First-Try" Rule: Why Traditional Diffusion Models Fail in Twitter Communities
A data-driven study of influences in Twitter communities
The paper presents a large-scale data-driven study of Twitter communities, analyzing 20.5 million profiles and 105 million tweets to understand user influence. The authors propose the First Influencer (FI) information diffusion model, which accounts for Twitter's interface-level suppression of duplicate messages, outperforming the traditional Independent Cascade (IC) model in prediction accuracy.
TL;DR
Researchers have long used the Independent Cascade (IC) model to predict how "viral" a tweet will go. This paper argues that IC is fundamentally flawed because it ignores a simple UI reality: Twitter hides duplicate tweets. The authors propose the First Influencer (FI) model, which recognizes that if the first person you follow who posts a link doesn't convince you to retweet it, you likely won't see it again from others. Their study of 20.5 million users proves this "first-take" approach is far more accurate for predicting social spread.
Background: Beyond the "Global" Twitter
While early studies treated Twitter as one giant monolith, this paper dives into Hashtag Communities (#android, #ladygaga, #marketing). The authors discovered that these niches behave differently—for instance, the #marketing community has a much higher average influence score than gamers in #android, and reciprocity (following back) is far higher within these interest-based clusters than in the general population.
The Core Problem: The Myth of Repeated Exposure
In classic cascade models, every time a new friend retweets a story, you get a fresh "roll of the dice" to be influenced.
The Reality Check: Twitter's interface suppresses duplicates. If three people you follow retweet the same news, it usually appears once in your timeline associated with the first person who shared it (or the most recent, depending on the algorithm version at the time). If you scroll past it once, the "influence attempt" is effectively over. This creates a "First Influencer" bottleneck that the IC model fails to capture.
Methodology: The First Influencer (FI) Model
The authors define a new state for users: Insusceptible.
- Selection: When a message spreads, the first person (the "First Influencer") to reach an inactive user makes an attempt with probability .
- Outcome: If the user doesn't retweet ( is not activated), they become insusceptible to that specific message.
- Resistance: Future attempts from other friends for the same message are set to a probability of zero ().
Table 1: Characteristics of the crawled Twitter communities, showing varying densities and interaction levels.
The Mathematical Twist: Non-Submodularity
Interestingly, by making the order of influence matter, the spread function loses its submodularity and monotony. In simple terms: adding more "seed" influencers doesn't always guarantee a larger spread in a predictable, diminishing-returns way, because influencers can now "compete" for the first-contact slot.
Experiments and Results
The authors compared the FI and IC models across seven distinct datasets.
1. Stability
A "good" model should be consistent. By splitting cascade logs into two sets, the authors found that the FI model had a significantly lower Root Mean Square Error (RMSE) in its predicted influence probabilities compared to the IC model. This suggests that the FI model captures an underlying truth about the data that the IC model misses.
2. Prediction Accuracy
When predicting the total number of people influenced by a single user (influence spread), the FI model consistently outperformed the IC model.
Figure 7: The FI model (red) tracks the ground truth of message spread much more closely than the IC model, which tends to overestimate reach in dense networks.
Critical Insight: Influence Homophily
The study also explored Homophily—the idea that "birds of a feather flock together." By taking the average of Klout and PeerIndex scores to create a "Digital Influence" (DI) score, they found that users tend to form mutual follow relationships with others of similar status. This suggests that influence isn't just a top-down broadcast; it's a horizontal reinforcement among peers.
Conclusion & Limitations
The First Influencer model represents a paradigm shift from theoretical "activation" to platform-aware "interaction." While the paper effectively proves that the first contact is the most vital, it assumes (zero influence after the first try). Future work might explore if there is a "reminder effect" where the second or third exposure has a non-zero, albeit much smaller, chance of influencing a user.
Takeaway for Marketers: In the age of algorithmic suppression, being the first to break a story to a community is far more valuable than being the tenth, even if you have more followers.
