Gender Identification in Greek Tweets: A Machine Learning Benchmark
A machine learning approach for gender identification of Greek tweet authors
This paper presents a machine learning framework for the gender identification of Greek Twitter authors using a newly curated corpus. The researchers utilize traditional ML algorithms like SVM and Multinomial Naive Bayes combined with TF-IDF encoding, achieving a peak accuracy of 0.70.
Executive Summary
TL;DR: This study addresses the challenge of Author Profiling—specifically gender identification—within the unique linguistic landscape of the Greek Twitter sphere. By collecting a new, randomized dataset of over 45,000 tweets and applying a rigorous preprocessing pipeline to manage Greek's complex morphology, the authors demonstrate that a Linear SVM with TF-IDF encoding can reach an accuracy of 70%. This work serves as a foundational baseline for Greek social media analytics.
Field Positioning: This is a practical, empirical benchmark study. While not introducing a new algorithm, it fills a critical gap in localized NLP by moving away from "celebrity-only" datasets to a "random population" model, providing a more realistic assessment of ML performance in Greek text classification.
The "Greek Problem": Motivation & Linguistic Hurdles
Why is gender identification harder in Greek than in English? The authors identify three primary "pain points":
- Stop Word Complexity: Unlike English, Greek stop words vary wildly due to tonal signs and the transition from Katharevousa (archaic) to Standard Modern Greek.
- Dataset Bias: Prior Greek studies often focused on a handful of famous authors (high-profile accounts), which does not reflect how the average user writes.
- Data Scarcity: Twitter does not provide gender metadata, making ground-truth annotation a manual and labor-intensive process.
To solve this, the authors created a "super list" of 2,175 stop words and manually annotated 463 distinct users to ensure a high-fidelity ground truth.
Methodology: From Raw Tweets to Vectors
The pipeline follows a classic Natural Language Processing (NLP) architecture:
- Data Acquisition: 500 random keywords were used to pull tweets, narrowing down to 463 validated authors.
- Preprocessing: Normalizing the polytonic/monotonic variations and applying the comprehensive stop-word filter.
- Vectorization: Using Bag-of-Words (BoW). Specifically, comparing simple Count Encoders (word frequency) against TF-IDF (Term Frequency-Inverse Document Frequency), which weights words by their uniqueness across the corpus.
Table 1: Performance of various classifiers across different encodings.
Experiments and Key Findings
The study compared five algorithms: Multinomial Naive Bayes, SVM (Linear & RBF), KNN, and Decision Trees.
- The Winner: The Linear Support Vector Machine (SVM) using TF-IDF encoding outperformed its peers, specifically in imbalanced scenarios (70% accuracy).
- The Role of Balancing: When the dataset was balanced (removing male instances to match female count), Naive Bayes showed the most stability, maintaining a 69% accuracy.
- Bias Insight: The authors noted that on imbalanced data, models were more likely to misclassify female instances, a common bias in social media datasets where male authors might exhibit more recognizable linguistic patterns in certain keyword clusters.
Figure 1: Comparison of misclassification rates between Naive Bayes and SVM on balanced vs. imbalanced data.
Critical Analysis & Future Outlook
Takeaway: The study proves that even without deep learning, traditional ML can extract significant gender-based stylistic signals from Greek text.
Limitations:
- Context Window: 70% accuracy is a solid start but remains insufficient for high-stakes applications like forensic linguistics.
- Handcrafted Features: The study relies on BoW, which ignores the syntactic structure that might be vital in a declension-heavy language like Greek.
Future Work: The authors suggest moving toward Neural Networks (CNNs/RNNs) and incorporating multi-modal data (e.g., profile images), which usually yields a 5-10% boost in similar English-language tasks.
