What Emotions Make One or Five Stars? Unmasking Model Bias through XAI
What Emotions Make One or Five Stars? Understanding Ratings of Online Product Reviews by Sentiment Analysis and XAI
The paper investigates the relationship between emotional sentiments in online reviews and star ratings using machine learning. It employs Explainable AI (XAI) techniques to analyze models like Random Forest and XGBoost, identifying "Joy" and "Negative Valence" as the primary predictors for Amazon product ratings.
TL;DR
This study moves beyond simply predicting star ratings from review text. While Random Forest models can accurately guess a product's rating using sentiment features like "Joy" or "Anger," the author uses Explainable AI (XAI) to prove that these models are often "right for the wrong reasons." The research uncovers that severe class imbalance in Amazon reviews (too many 5-star ratings) leads models to develop flawed logic that can only be caught through local diagnostic tools.
Background: The Black Box of Consumer Sentiment
In the world of e-commerce, reviews are gold. Most research focuses on Accuracy: "Can we predict a 5-star rating based on text?" However, this paper asks a more critical question: "Why did the model think this was 5 stars?" By leveraging NLP to extract emotions (based on Ekman’s universal emotion theory) and emotional valence (positive/negative), the study attempts to bridge the gap between human feeling and machine prediction.
Methodology: Auditing the Machine
The research followed a three-step workflow:
- Benchmarking: Testing algorithms like KNN, SVM, Random Forest, and XGBoost.
- Global Analysis: Using Feature Importance to see which emotions matter most across the whole dataset.
- Local XAI: Using Local Feature Attributions and Partial Dependency Plots (PDP) to see how the model behaves for a single specific review.
Figure 1: Benchmarking shows Random Forest (rf) achieving the lowest RMSE, but this is only half the story.
The "Aha!" Moment: When Logic Breaks
The Global Feature Importance (Figure 2) suggested that Joy and Negative Valence were the strongest predictors. On the surface, this makes sense. But when the author looked at individual cases using Local Attributions, the model's "thinking" appeared erratic.

For instance, in some cases, a high score for positive valence actually decreased the predicted rating. This is a logical contradiction. The Partial Dependency Plots (Figure 4) further revealed that while "Anger" and "Fear" were correctly linked to lower ratings, the relationships were surprisingly weak, and "Positive Valence" often had a zero slope—meaning the model was effectively ignoring it.
Figure 3: Local attribution reveals that a high intercept (base rating) dominates the prediction, masking the actual influence of sentiments.
The Culprit: Dataset Bias
Why would a high-performing model have such flawed logic? Study 3 provides the answer. By reframing the task as a classification problem, the author found a No-Information Rate of 64.4%.
In plain English: because the vast majority of Amazon reviews are 5-star ratings, a model can achieve ~70% accuracy just by "guessing" 5 stars most of the time. The model wasn't learning the nuances of human emotion; it was learning the statistical imbalance of the dataset.
Critical Insight & Conclusion
This paper serves as a vital warning for data scientists in the NLP and RecSys space:
- Performance != Understanding: A low RMSE or high Accuracy does not mean your model has "solved" sentiment.
- XAI as a Diagnostic: Tools like PDP and Local Attributions are not just for "explaining" results to stakeholders; they are essential for debugging and identifying when a model is leaning on dataset bias (shortcuts) rather than features.
- The Reality of Reviews: Consumer review data is heavily skewed. Without addressing class imbalance, sentiment analysis remains a game of predicting the majority class.
For future work, the study suggests that we must move beyond basic emotion lexicons and utilize XAI to refine feature sets—discarding features with low importance or illogical local variances to build more robust, fair, and interpretable systems.
