Decoding Deception: Deep Syntactic Fingerprinting for Age and Gender Prediction

Age and Gender prediction in Open Domain Text

2020-01-01
Emad E. Abdallah, Jamil R. Alzghoul, Muath Alzghool
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an automated framework for predicting the age and gender of authors in open-domain deceptive text using machine learning. By leveraging a unique combination of Unigrams, Part-of-Speech (POS), and Context-Free Grammar (CFG) production rules, the authors achieved a state-of-the-art accuracy of 82.81% for gender and 83.2% for age prediction.

TL;DR

As social media becomes the primary medium for interaction, demographic deception (lying about age or sex) has surged. This paper introduces a machine learning approach that looks past what people say to how they structure their sentences. By using Context-Free Grammar (CFG) production rules and Support Vector Machines (SVM), the authors pushed prediction accuracy for both age and gender to over 82%, a massive leap over previous baselines.

Problem & Motivation: The Mask of the Open Domain

In "Open Domain" text—think random forum posts or dating site messages—there is no fixed context. This makes it incredibly difficult for standard algorithms to tell if a "19-year-old female" is actually a middle-aged male.

Prior works often relied on Unigrams (individual words). However, vocabulary is easily faked. The authors' insight was that while a deceiver can change their words, they rarely change their syntactic habits—the subconscious way they produce sentence structures. These "Production Rules" serve as a deeper, more resilient fingerprint.

Methodology: Beyond Simple Word Counts

The researchers developed a rigorous pipeline using the Weka machine learning framework. The core innovation lies in the selection and combination of feature sets:

  1. Unigrams: Captures vocabulary.
  2. POS (Part of Speech): Captures grammatical categories (nouns, verbs, etc.).
  3. Production Rules (CFG): Captures the "deep syntax" or paths taken to build a sentence.

The Pipeline

The process follows a six-step journey: Input -> Tokenization -> String to Word Vector (cleaning) -> Feature Selection (Information Gain) -> Classification (SVM, Naïve Bayes, etc.) -> Evaluation.

Methodology Overview Fig 1: The proposed methodology for age and gender detection.

Experiments & Results: A New SOTA

The authors tested several classifiers, including single models (SVM, Decision Trees) and ensemble methods (Meta classifiers).

1. Gender Prediction

The results were striking. Before applying feature selection, accuracy hovered around 69%. However, once Information Gain was used to filter out noise, the SVM using CFG features reached 82.81%. This validates that structural rules are far more predictive of gender than vocabulary alone.

2. Age Prediction

Age prediction is historically harder than gender. The study split participants into "Younger" (≤35) and "Elder" (>35). By combining POS and Lexicalized Production Rules with ensemble methods, they achieved 83.20% accuracy.

Accuracy vs Features Fig 2: Impact of different feature sets on prediction accuracy after feature selection.

Critical Analysis & Conclusion

Why it Works

The success of this approach lies in the Inductive Bias that syntax is a more stable demographic marker than semantics. In open-domain text, where the topic changes constantly, lexical features (words) are too "noisy." Syntactic units (how a person transitions from a Verb Phrase to a Noun Phrase) are less likely to be modified even when a person is trying to be deceptive.

Limitations & Future Work

While the accuracy is high, the binary split for age (under/over 35) is somewhat broad. Real-world applications might require more granular age "brackets" (e.g., teenagers vs. seniors). Additionally, as LLMs (Large Language Models) become more prevalent, future research will need to determine if these "syntactic fingerprints" still exist in AI-generated deceptive text.

Final Takeaway

This work proves that deep syntax is the key to unmasking online deception. For cybersecurity and social media platforms, integrating CFG-based analysis could be a powerful tool in verifying user identity and protecting vulnerable populations from deceptive actors.

Find Similar Papers

Try Our Examples

  • Search for recent studies that use Deep Learning or Transformers (like BERT) for age and gender prediction in deceptive social media contexts.
  • Which paper first introduced the use of Context-Free Grammar (CFG) production rules for stylometry, and how has its application evolved in deception detection?
  • Investigate how demographic prediction methods from open-domain text are being applied to identify "social bots" or "sockpuppet" accounts on platforms like X (Twitter) or Reddit.
Contents
Decoding Deception: Deep Syntactic Fingerprinting for Age and Gender Prediction
1. TL;DR
2. Problem & Motivation: The Mask of the Open Domain
3. Methodology: Beyond Simple Word Counts
3.1. The Pipeline
4. Experiments & Results: A New SOTA
4.1. 1. Gender Prediction
4.2. 2. Age Prediction
5. Critical Analysis & Conclusion
5.1. Why it Works
5.2. Limitations & Future Work
5.3. Final Takeaway