Beyond Style: Identifying Social Media Authors Through Semantic DNA

Content-Based Authorship Identification for Short Texts in Social Media Networks

2021-01-01
José Gaviria de la Puerta, Iker Pastor-López, Javier Salcedo-Hernández, Alberto Tellaeche, Borja Sanz, Hugo Sanjurjo-González, Alfredo Cuzzocrea, Pablo García Bringas
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a content-based authorship identification method specifically for short social media texts (tweets). By utilizing Word2Vec skip-gram embeddings and a vector-averaging approach, the system identifies authors based on semantic topics rather than traditional stylistic traits, achieving a top-1 accuracy of 87.5% across a heterogeneous multi-language dataset.

TL;DR

Researchers have developed a new way to unmask anonymous authors on social media by analyzing what they talk about rather than just how they write. By using Word2Vec embeddings to create "topic fingerprints," this approach achieves 87.5% accuracy in identifying users across short, restricted texts like tweets, even when traditional stylistic analysis fails.

The Problem: When Style is Too Short to See

In the world of forensic linguistics, Authorship Identification has traditionally been about "Stylometry"—the study of linguistic style (e.g., how often one uses semicolons or specific function words). However, on platforms like Twitter (X), the character limit creates a "sparsity problem." There simply isn't enough text in a single tweet to build a reliable stylistic profile.

When dealing with malicious actors—those spreading radicalization or defamation via fake profiles—investigators need a way to link these "trolls" to their real identities. Current methods are often blind if the offender avoids their usual writing quirks.

The Insight: Content is Recurrent

The authors of this paper pivot from Style to Content. Their core hypothesis is simple: People are creatures of habit. A user who writes about right-wing politics, local football, and technological gadgets on their real profile is highly likely to use similar vocabulary and discuss the same topics on their fake profile.

Methodology: Mapping the Semantic Manifold

The team constructed a sophisticated pipeline to transform messy social media data into a mathematical coordinate system:

1. The Embedding Layer

They used Word2Vec (Skip-gram model). Unlike older methods like TF-IDF, Word2Vec captures the contextual meaning of words. If User A talks about "automobiles" and User B talks about "cars," Word2Vec recognizes these as semantically adjacent.

2. Tweet & User Mean Vectors

To represent an entire tweet, the authors calculated the Mean Vector of all its words. For a user profile, they averaged all the tweet vectors. This creates a "centroid" in a 400-dimensional space that represents the user's "Topic DNA."

Proposed Approach Logic

3. Classification via Proximity

To identify an author, they employed a k-Nearest Neighbors (k-NN) algorithm. Instead of Euclidean distance, they used Cosine Similarity to measure the angle between vectors, effectively asking: "Which known author's topic cloud does this new tweet most closely align with?"

Experimental Results

The study utilized a diverse dataset of 142,000 tweets from 40 heterogeneous users across multiple languages (Spanish, Catalan, Basque) and topics (Politics, Sports, Humor).

MetricResult
General Accuracy (Top-1)87.5%
Optimistic Accuracy (Top-3)97.5%

Even with users who had a low number of publications (like User 26), the system successfully identified them by linking their sparse vocabulary to broader semantic themes.

User Distribution and Dataset Overview

Critical Insights & Future Horizon

The success of this content-based approach suggests that in constrained environments (like OSNs), Semantics > Syntax.

Key Strengths:

  • Language Agnostic: The methodology works across different languages by training on multi-lingual corpora.
  • Candidate Reduction: Even if it doesn't get the author right on the first try, it reduces the list of suspects to a tiny, highly-relevant pool (97.5% accuracy in the Top-3).

Limitations & Future Work: The authors acknowledge that simple vector averaging is a "baseline" compared to more modern architectures like Paragraph Vectors or Recursive Neural Networks. Future iterations could incorporate these to capture the order of words and nuances in sentiment, potentially pushing the accuracy even closer to 100%.

Conclusion

This research provides a vital tool for digital forensics and counter-terrorism. By proving that our interests and vocabulary serve as a "semantic fingerprint," the study opens up a new frontier in keeping social networks accountable and safe.

Find Similar Papers

Try Our Examples

  • Find recent papers that combine semantic word embeddings with stylistic features (Stylo-Semantic) for cross-platform authorship verification.
  • Which paper first introduced the Paragraph Vector (Doc2Vec) approach, and how does it specifically outperform the vector-averaging method used in this study for short-text classification?
  • Explore newer research applying Contrastive Learning or Transformer-based embeddings (like BERT) to the task of identifying "troll" or radicalized profiles in social media.
Contents
Beyond Style: Identifying Social Media Authors Through Semantic DNA
1. TL;DR
2. The Problem: When Style is Too Short to See
3. The Insight: Content is Recurrent
4. Methodology: Mapping the Semantic Manifold
4.1. 1. The Embedding Layer
4.2. 2. Tweet & User Mean Vectors
4.3. 3. Classification via Proximity
5. Experimental Results
6. Critical Insights & Future Horizon
7. Conclusion