Deciphering the Digital Ballot: Why Your Nouns Reveal Your Politics

Content-based Classification of Political Inclinations of Twitter Users

2018-12-01
Marco Di Giovanni, Marco Brambilla, Stefano Ceri, Florian Daniel, Giorgia Ramponi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a NLP-based framework to classify the political inclinations of Italian Twitter users by analyzing the syntactic features of their tweets. By focusing specifically on Italian deputies as a ground truth, the authors demonstrate that a Multinomial Logistic Regression model using TF-IDF vectorized nouns can achieve a SOTA-level classification accuracy of 89%.

TL;DR

Can your choice of words summarize your political soul? Researchers at Politecnico di Milano have proven that for Italian politicians on Twitter, the answer is a resounding "Yes." By stripping away social connections and focus purely on syntactic features, they achieved an 89% accuracy in predicting political party affiliation, discovering that the "secret sauce" of political identity lies almost exclusively in the nouns we choose to use.

The Motivation: Beyond the Social Graph

Traditional social media analysis often treats political leanings as a "contagion" moving through a network—if your friends are right-wing, you likely are too. However, this Network Analysis approach has two flaws: 1) It requires access to private connection data, and 2) It ignores the actual message being communicated.

The authors of this paper pivoted to a content-centric hypothesis: Political identity is a dialect. They set out to find if specific parts of speech—nouns, verbs, or adjectives—hold the strongest "fingerprint" of a party’s ideology.

Methodology: The Anatomy of a Tweet

The researchers focused on four major Italian parties: Movimento 5 Stelle, Lega, Partito Democratico, and Forza Italia. Their pipeline involved:

  1. Isolation: Extracting tweets from 188 verified deputies to establish an ironclad ground truth.
  2. Filtering: Using POS (Part-of-Speech) tagging to separate the text into lists of nouns, verbs, adjectives, and adverbs.
  3. Vectorization: Converting these lists into numerical vectors using TF-IDF (Term Frequency-Inverse Document Frequency). This mathematical trick rewards words that are unique to a specific politician/party while penalizing common "filler" words.

Model Pipeline and Process Note: The process involves transforming raw tweets into lemmatized syntactic categories before feeding them into a Multinomial Logistic Regression model.

The "Noun" Breakthrough

The most striking finding was the disparity between different parts of speech. When the model only looked at adverbs, it was barely better than a coin flip (50% accuracy). When it looked at verbs, it struggled.

However, Nouns were the gold mine.

  • Accuracy with Nouns: 0.89
  • Accuracy with Every Word: 0.86

This suggests that including "style" elements (how people describe things or actions) actually adds noise rather than signal. Local political identity is defined by the entities we talk about (e.g., "citizen," "pension," "mosque") rather than the way we describe them.

Confusion Matrix of Political Prediction Fig 1: The confusion matrix reveals that while most parties are highly distinct, "Lega" members are sometimes misclassified as "M5S," suggesting a syntactic overlap in populist vocabulary.

Deep Insights: The Semantic Map

By using t-SNE to project high-dimensional word vectors into a 2D space, the researchers visualized the "cohesion" of political parties.

  • Cohesive Parties: The Partito Democratico and M5S showed tight clusters, meaning their members stay "on script."
  • Fragmented Parties: Lega members were more spread out, indicating a less centralized or more diverse vocabulary.

The study even identified a "lexical tug-of-war." For instance, the word "citizen" (cittadino) is a high-weight noun for populist parties, while institutional parties favor words like "minister" or "commitment."

Critical Analysis & Future Work

While the 89% accuracy is impressive, the study has limitations. It was performed on professional politicians who are trained to stay "on message." The real challenge—and the authors' next goal—is applying this to the general public. Does a regular voter's vocabulary mirror their chosen representative's nouns well enough to predict their vote?

Furthermore, the study highlights the "Lega" ambiguity, where a lack of a specific "dictionary" makes classification harder. This suggests that some political movements are defined more by what they oppose than the specific labels they use.

Conclusion

This research proves that political inclination isn't just about who you follow; it's about the very building blocks of your sentences. In the age of AI, our "noun-print" may be just as identifying as a fingerprint.

Find Similar Papers

Try Our Examples

  • Search for recent papers that combine content-based classification with Graph Neural Networks (GNNs) for predicting political orientation on Twitter.
  • Which study first identified the "nouns-over-verbs" preference in political linguistics, and how does this paper's TF-IDF approach quantify that preference?
  • Find research applying these syntactic classification techniques to detect political polarization or shift in non-Western democratic social media contexts.
Contents
Deciphering the Digital Ballot: Why Your Nouns Reveal Your Politics
1. TL;DR
2. The Motivation: Beyond the Social Graph
3. Methodology: The Anatomy of a Tweet
4. The "Noun" Breakthrough
5. Deep Insights: The Semantic Map
6. Critical Analysis & Future Work
6.1. Conclusion