Multilingual Pornography Detection on Twitter: A Comparative Machine Learning Approach
Twitter Pornography Multilingual Content Identification Based on Machine Learning
2017-01-01
Summary
Problem
Method
Results
Takeaways
Abstract
The paper explores "Twitter Pornography Multilingual Content Identification" using machine learning techniques (Decision Tree, Naive Bayes, and SVM). It focuses on identifying pornographic vs. non-pornographic tweets in Indonesian, English, and a combination of both, utilizing TF-IDF for feature extraction.
## TL;DR
This research addresses the critical challenge of identifying pornographic content on Twitter across multiple languages (Indonesian and English). By benchmarking **Decision Trees**, **Naive Bayes**, and **Support Vector Machines (SVM)** using a TF-IDF framework, the study reveals that while Naive Bayes is highly effective for Indonesian text (92.78% accuracy), SVM with a Linear Kernel is superior for handling the complexities of English and mixed-language environments.
## Problem & Motivation
The proliferation of adult content on social media presents a significant risk to the morale and safety of minors. Prior works often relied on static **URL Blocking** or simple **Keyword Filtering**. However, Twitter poses a unique challenge:
* **Dynamic Language**: Natural language on Twitter is irregular, filled with slang, and frequently skips standard grammar.
* **Multilingualism**: Users often switch between Indonesian and English, complicating traditional monolingual filters.
* **Evasion**: Content creators use creative text variations to bypass simple keyword-based blocks.
The authors' insight was to move beyond simple filters toward a **Machine Learning classification** approach that analyzes the statistical significance of words (Unigrams) within documents to differentiate between "Positive" (Safe) and "Negative" (Pornographic) content.
## Methodology
The system follows a classic NLP pipeline optimized for the noisy nature of social media data:
1. **Pre-processing**: This involves aggressive cleaning (removing links, hashtags, and special characters) followed by **Porter Stemming** (adapted for Indonesian) to reduce words to their base forms.
2. **Feature Engineering**: The team utilized **TF-IDF (Term Frequency-Inverse Document Frequency)** to weight the importance of words. They focused on **Unigrams** (single-word features) like "sex" or "porn."
3. **Classification**: Three main algorithms were compared:
* **Decision Tree**: A flowchart-like structure for attribute testing.
* **Naive Bayes**: A probabilistic classifier based on the assumption of independence between variables.
* **Support Vector Machines (SVM)**: A high-dimensional mapping technique aimed at finding the optimal hyperplane for classification.

## Experiments & Results
The researchers used a dataset of 600 tweets (equally split between pornographic and non-pornographic across three categories: Indonesian, English, and Combined).
### Performance Breakdown
The results highlighted a fascinating discrepancy between language performance:
* **Indonesian Dataset**: **Naive Bayes** dominated with an average accuracy of **92.78%**. The authors noted that Indonesian grammar, while complex, showed high regularity in the specific pornographic corpus used.
* **English Dataset**: **SVM** took the lead with **83.33%**.
* **Combined Dataset**: Performance dropped to **72.72%** (SVM), illustrating the difficulty of "code-switching" where users mix languages in a single post.

### The Kernel Selection (SVM)
A deep dive into SVM kernels for the Indonesian dataset proved that the **Linear Kernel** is most effective for text classification (87.60% average). More complex kernels like Radial or Polynomial actually degraded performance, suggesting that text features are often linearly separable in a high-dimensional TF-IDF space.
## Critical Analysis & Conclusion
### Takeaway
The study confirms that a "one-size-fits-all" algorithm is rarely optimal for multilingual NLP. **Naive Bayes** remains a powerful, lightweight tool for specific linguistic structures (like Indonesian), whereas **SVM** provides the necessary robustness for the more fragmented English-speaking "Twitter-verse."
### Limitations
1. **Dataset Size**: 600 tweets is relatively small for a modern machine learning task; larger datasets would likely reveal more edge-case failures.
2. **Context**: Unigrams cannot capture sarcasm or nuanced meanings that depend on word order (Bigrams/Trigrams).
### Future Outlook
With the rise of Large Language Models (LLMs), the next step for this research would be to apply **transformers** to capture the semantic context of tweets, potentially closing the gap in the "Multilingual Combined" category where accuracy currently struggles.
