Predicting eWOM Success: Why Content and Author Synergy is the Key to Viral Advertising
Predicting the influence of users’ posted information for eWOM advertising in social networks
This paper introduces a predictive framework for estimating the influence of social network posts to optimize eWOM (electronic Word-of-Mouth) advertising. It proposes two main models—multiple regression and an ensemble classification system—to predict the "influence score" of a post by integrating both author-related and post-content features.
TL;DR
In the world of social media advertising, not all "likes" are created equal. This paper presents a sophisticated framework to predict the influence of Facebook posts by combining the DNA of the content (sentiment, media type) with the social weight of the author. By moving beyond simple counts to a weighted "Influence Score," the authors achieved a major boost in prediction accuracy using an ensemble machine learning approach.
Background: The eWOM Dilemma
Electronic Word-of-Mouth (eWOM) is the holy grail of modern marketing. Platforms like Facebook allow marketers to turn user posts into "sponsored stories," but picking the right post is a gamble. Historically, researchers looked at either who is important (the influencer) or what makes a post popular (the content), but rarely both. This paper bridges that gap, arguing that a post's true power is a product of its creator's reach AND its intrinsic quality.
The "Influence Score": A Smarter Metric
The authors identify a critical flaw in current metrics: treating every response as equal. They propose a weighted Influence Score:
- Physical Intuition: A "like" from a user with 5,000 friends carries more weight than one from a user with 50.
- Formula: The weight of a user is defined as their friend count normalized by the maximum friend count in the network: The total influence of a post is the sum of the weights of all users who interacted with it.
Methodology: The Ensemble Engine
The research moves beyond simple linear regressions (which struggle with complex, non-linear social behaviors) to an Ensemble Model.
1. Feature Engineering
The authors categorize predictors into:
- Content: Length, sentiment (positive/negative ratio), and media type (links/photos).
- Authorial: Friend count, but also novel features like
avg_like(the author's historical "gravity"). - Temporal: When the post was made (weekday vs. weekend).
2. The Model Architecture
The system utilizes a "Committee of Experts" approach. By aggregating predictions from five diverse algorithms, the model mitigates the weaknesses of any single approach:

Experimental Results
The authors tested their models on a dataset of 510 Facebook posts. The results were clear:
- Weight Matters: Influence score and simple "like counts" are correlated but not identical (see Figure 3). Relying on counts alone leads to significant marketing errors.
- Performance Breakthrough: Using their specialized feature set (Set 3), they achieved a Macro-average F1-score of 0.62-0.65, nearly doubling the performance of models using only standard literary features.
Fig 3: The disparity between "Influence Score" and "Like Count" justifies the need for weighted metrics.
Critical Insight: The Sentiment Paradox
Interestingly, the regression analysis found that the number of negative words and the ratio of negative words were highly significant predictors of influence. This suggests that in social environments, strong sentiment (even if negative) can drive higher engagement and "reach" than neutral content—a vital takeaway for marketers looking to spark conversation.
Conclusion & Limitations
This work demonstrates that predicting social influence is a multi-modal problem. By combining author history with real-time content sentiment, marketers can move from "guessing" to "calculating" virality.
Limitations to consider:
- The sample size (510 posts) is relatively small for modern deep learning standards.
- The study is focused on Facebook; dynamics on TikTok or X (Twitter) may vary due to algorithm differences.
Future research should look toward Graph Neural Networks to automatically extract the "Topological" features that this paper had to manually engineer.
