Unmasking the Redditor: Advanced Authorship Attribution in Forum Ecosystems
Authorship Attribution using data from Reddit forum
This paper investigates Authorship Attribution (AA) on the Reddit forum using comments from the "r/science" subreddit. It compares multiple text representations (n-grams of words, characters, and POS tags) across various classification scenarios using online learning algorithms, achieving high accuracy (often >95%) in identifying authors.
TL;DR
Researchers from the University of São Paulo analyzed over 680,000 comments from Reddit's "r/science" to determine if an author's unique "digital fingerprint" can be identified amidst thousands of users. By leveraging batch-based online learning and multi-level n-gram analysis, the study achieved near-perfect accuracy, proving that how you "decorate" your text—spacing, punctuation, and formatting—is just as revealing as the words you choose.
Context & Positioning
In the era of information warfare, Authorship Attribution (AA) has moved from historical literary analysis (like solving the mystery of The Federalist Papers) to a critical tool for fighting Fake News and Astroturfing. This paper, published at the Brazilian Symposium on Information Systems (SBSI'20), positions itself as a scalable solution for forum-style social media, where text is longer and more formatted than Twitter but more chaotic than journalism.
The Problem: The Noise of the Crowd
Most AA methods fail on social media for two reasons:
- Computational Complexity: Processing millions of comments using traditional "offline" models exhausts memory.
- Informal Syntax: Slang, abbreviations, and typos break standard linguistic models.
The authors argue that these "errors" and formatting choices (bolding, italics, tabulations) are not noise—they are features.
Methodology: Capturing the Stylistic Fingerprint
The core of the approach involves three distinct feature representations:
- Character n-grams: Captures the use of special characters, emoticons, and white space.
- Word n-grams: Captures vocabulary and common phrases.
- POS Tagging: Captures the underlying grammatical structure (syntax).
The research utilized three high-scale classifiers: Stochastic Gradient Descent (SGD), Perceptron, and Passive-Aggressive. The latter was particularly effective because it uses a regularization constant () to ignore outliers, making it robust against the erratic nature of Reddit posts.
Figure: ROC Curve of the Perceptron classifier for binary authorship identification.
Experiments and Key Findings
The researchers tested three scenarios:
- Binary: Distinguishing between the two most active users.
- Multiclass: Distinguishing between the top 10 most active users.
- One-vs-All: Identifying one specific author out of the entire subreddit.
The "Word vs. Character" Inversion
A fascinating discovery was the behavior of in different contexts:
- Word n-grams peaked at . When increased further, the model became "too specific" (overfitting), and accuracy dropped.
- Character/POS n-grams required or . Because characters carry less information individually, the model needed longer sequences to "see" the author’s style.
Table: Experimental results showing Precision (P), Recall (R), and F1-Score for the top 10 authors.
Deep Insight: Why Passive-Aggressive Won
While SGD is a standard industry workhorse, it is sensitive to learning rates and hyperparameter tuning. In this study, the Passive-Aggressive classifier was the MVP. It remains "passive" when a classification is correct and becomes "aggressive" only when an error occurs, adjusting the model just enough to fix the mistake without overreacting to outliers. This makes it ideal for the "wild west" of Reddit comments.
Critical Analysis & Future Outlook
Takeaway: This work demonstrates that even in a sea of millions of users, our writing habits—down to how we use bold text or place our commas—are remarkably unique.
Limitations: The study focuses on "high-volume" authors. Identifying "casual" users who only post a few sentences remains a significant "cold-start" challenge in the AA field.
Next Steps: The authors suggest incorporating Sentiment Analysis to see if an author's emotional volatility is also a consistent identifier. This could lead to even more resilient systems for detecting botnets and coordinated disinformation campaigns.
Disclaimer: This analysis is based on the research paper "Authorship Attribution using data from Reddit forum" by Casimiro and Digiampietri (2020).
