SentiMeter-Br: Mastering the Nuances of Brazilian Consumer Sentiment
SentiMeter-Br: A Social Web Analysis Tool to Discover Consumers' Sentiment
This paper introduces SentiMeter-Br, a specialized social web analysis tool designed to decipher consumer sentiment in Brazilian Portuguese. By employing a domain-specific dictionary and a custom sentiment calculation engine, it significantly out-performs general-purpose tools like SentiStrength in capturing local linguistic nuances.
TL;DR
SentiMeter-Br is a specialized analytical framework designed to bridge the gap in Brazilian Portuguese sentiment analysis. By moving beyond simple word lists to include grammatical rules for tenses and negations, it achieves a high correlation with human experts (0.89 Pearson). It also integrates a cross-platform framework for extracting data from Twitter and Facebook, providing a comprehensive tool for market research.
Background & Motivation: Why General Tools Fail
Sentiment analysis is often treated as a simple "bag of words" problem. However, in the Brazilian market, context is everything. A word that is positive in one domain might be negative in another. Furthermore, existing tools like SentiStrength often treat all words equally, ignoring the fact that "I love this soap" (present) carries more weight for a brand than "I loved this soap" (past).
The authors identified that generic dictionaries fail to capture:
- Domain-Specific Slang: Local expressions used in niche markets (e.g., hair cosmetics).
- Negation Flipping: Mathematical errors where "not bad" is incorrectly categorized as highly negative.
- Tense Relevance: The diminishing importance of past-tense experiences in real-time social media monitoring.
Methodology: The Logic of SentiMeter-Br
The system's architecture relies on a specialized PT-Br dictionary curated by specialists. The sentiment calculation isn't just a sum; it's a weighted division based on the linguistic structure.
1. The Sentiment Strength Formula
The core logic involves a formula where the sum of word polarities is divided by the square root of the total word count, adjusted by a "Tense Factor" ().
If a verb is in the past tense, the value increases, effectively lowering the overall sentiment "strength." This rewards "present-tense" feedback, which is more actionable for companies.
2. Handling Exceptions
Unlike standard models, SentiMeter-Br uses specific files for NEG-FILE (not, never) and NEG-ADJ-FILE (bad, ugly). When a negation precedes a negative adjective, the system applies an exception rule to flip the polarity toward a neutral or slightly positive value, mimicking human intuition.
Figure 1: The framework facilitates mobile access and multi-platform data harvesting.
Experiments & Results
The researchers validated the Dictionary using both specialist human review and Machine Learning (via the Weka software).
Machine Learning Validation
They tested multiple algorithms, including Naive Bayes and Decision Trees, but Sequential Minimal Optimization (SMO)—a type of Support Vector Machine—emerged as the winner.
| Algorithm | Positive F-Measure | Negative F-Measure |
|---|---|---|
| Decision Tree | 0.74 | 0.88 |
| Naive Bayes | 0.75 | 0.85 |
| SMO | 0.87 | 0.91 |
Interestingly, the researchers found that while bigrams and trigrams didn't significantly shift the needle, the removal of stopwords improved classification accuracy by over 18% in some configurations.
Table IV: Comparison of F-Measures across different classification algorithms.
Critical Insight & Conclusion
The success of SentiMeter-Br lies in its linguistic sensitivity. By acknowledging that Brazilian Portuguese speakers use "not bad" to mean "good" and that past experiences are less intense than current ones, the tool achieves a 0.89 Pearson correlation—strikingly close to human-level agreement.
Takeaway for Practitioners: When building sentiment engines for non-English markets, a "lexicon-plus-grammar" approach still holds significant value, especially in domain-specific contexts like cosmetics or fashion where specialized slang dominates.
Future Work: The authors aim to expand this into other sectors like education and technology, and further refine geographic tracking for Twitter data to match the granularity of their Facebook framework.
