Unmasking the Puppeteers: Identifying Stock Speculators and Influencers via Social Media Sentiment
Speculator and Influencer Evaluation in Stock Market by Using Social Media
Identifying speculators and influencers in the stock market using Twitter sentiment analysis and machine learning across 6 NASDAQ giants (AAPL, AMZN, GOOG, TSLA, MSFT). The study utilizes the Loughran-McDonald (LMD) financial dictionary and achieves a 58% prediction accuracy using RBF-Kernel SVM.
TL;DR
Social media is no longer just a place for "feelings"; it is a digital battleground for market manipulation. This paper investigates how to identify Speculators (those spreading disinformation for profit) and Influencers (voices that sway buying decisions) by correlating individual Twitter sentiment with the price movements of major companies like Tesla and Apple. Using an RBF-Kernel SVM and the Loughran-McDonald financial dictionary, researchers achieved 58% accuracy in linking specific users to market returns.
Problem & Motivation: The Noise in the Machine
Why is it so hard to predict the stock market using Twitter? Most researchers look at the volume of tweets. However, this study finds that the correlation between tweet count and trading volume is surprisingly weak (often below 0.5).
The real problem lies in distinguishing between general chatter and strategic manipulation. Traditional methods treat all users as equal, but in reality, a few "Influencers" or "Speculators" hold disproportionate power. Most sentiment analysis also fails because it uses standard dictionaries—where a word like "liability" might be negative in general conversation but neutral or purely descriptive in a financial context.
Methodology: Mapping Sentiment to Money
The researchers developed a pipeline to transition from raw text to market-aligned labels.
1. Financial Sentiment Mapping
Instead of generic sentiment tools, the study uses the Loughran-McDonald (LMD) dictionary, which categorizes 86,486 words into 8 dimensions: Positive, Negative, Uncertainty, Litigious, Constraining, Superfluous, Interesting, and Modal.
2. The Model Architecture
The team collected 3.7 million tweets over 5 years. They focused on "Close-to-Close" price returns () to capture overnight manipulative activity.
Fig 1: The architecture of the Speculator/Influencer identification process.
The crucial step is Threshold Optimization. By setting a threshold ( or 150 bps), the authors filtered out the 70% of "neutral" noise, focusing only on "extreme" tweets that actually correlate with significant price swings.
Experiments & Results: Why SVM Wins
The researchers tested 8 different machine learning methods. Interestingly, complex models like Multi-Layer Perceptrons (MLP) and AdaBoost struggled, hovering around 38% accuracy.
The RBF-Kernel SVM emerged as the clear winner with 58% accuracy. This suggests that the relationship between financial sentiment and stock returns is non-linear and operates best in a high-dimensional feature space.
Fig 2: The "Up-side-down Sine" relationship showing how increasing the movement threshold improves prediction accuracy by filtering noise.
Key Insight: The Speculator Signature
By analyzing the "effect rate" of individual users, the researchers identified a small group of people who consistently post extreme sentiment (highly positive or highly negative) that aligns with subsequent market movements. These are the "Possible Speculators."
Fig 3: Distribution of users by their influence effect on AAPL stock.
Critical Analysis & Conclusion
Takeaway
The study proves that individual-level analysis is more valuable than aggregate "public opinion" when trying to spot market manipulation. By narrowing the focus to high-impact users and using domain-specific dictionaries, we can begin to identify the sources of artificial market volatility.
Limitations
- The Identity Problem: The study acknowledges that identifying a "speculator" legally is difficult; they can only label users as "possible" influencers.
- Twitter's API/Selenium: Data collection was hampered by Twitter's rate limits, requiring Selenium-based scraping which is harder to scale.
- Spelling & Slang: Financial dictionaries often miss Twitter-specific slang (e.g., "$TSLA to the moon"), which could be solved with modern Transformer-based models (Future Work).
Future Outlook
The next frontier is Noise Reduction. If we can automatically filter out bots and general retail noise, the "signal" from true market manipulators will become even clearer, potentially providing a real-time warning system for investors against pump-and-dump schemes.
