Inferring Lurkers’ Gender: Solving the "Silent User" Problem via Interest Tags
Inferring Lurkers’ Gender by Their Interest Tags
This paper introduces a specialized framework for predicting the gender of "lurkers"—social media users who consume content without posting—using only their self-selected interest tags. By employing a "Conceptual Class" expansion method based on association rule mining, the authors successfully condense sparse tag spaces to improve classification accuracy on Sina Weibo data.
TL;DR
How do you identify the gender of a social media user who never posts a single word? Researchers from Wuhan University have addressed this "lurker" problem by shifting the focus from what users say (text) to who they want to be (interest tags). By grouping sparse, diverse tags into "Conceptual Classes" and expanding them via association mining, they achieved a significant boost in classification accuracy on Sina Weibo.
The "Lurker" Motivation: Beyond the Vocal Minority
Most social media research suffers from "participation bias"—we study the people who talk. However, in platforms like Sina Weibo, nearly 7.4% of the population are lurkers: users who browse but never post.
Current SOTA methods for gender detection rely on:
- Lexical features: N-grams and vocabulary.
- Stylistic features: Punctuation usage and slang.
- Syntactic features: Sentence structure.
For a lurker, these features are non-existent. The only clue left is the Interest Tags in their profile. But tags are a nightmare for machine learning: they are sparse (usually <6 per person) and highly idiosyncratic (e.g., two fans of different F1 teams might not share a single tag).
Methodology: The Power of Conceptual Classes
The core insight of this paper is that while tags are diverse, the underlying concepts are gender-linked. The authors propose a "Conceptual Class" (CC) framework to condense the sparse tag space.
1. Building the Foundation
They defined 27 conceptual classes based on social and psycholinguistic traits. For instance:
- Female-oriented: Beauty, Cosmetic, Family, Fashion.
- Male-oriented: Cars, Technology, Sports(M), Politics.
2. Expanding the Vocabulary (ECC)
To handle the "Long Tail" of tags, the authors developed an Expanding Conceptual Class (ECC) algorithm. Using the Apriori algorithm on a massive unlabeled dataset, they looked for tags that frequently co-occur with the "seed" tags in their conceptual classes.
Figure 1: Examples of initial vs. expanded tags. Notice how specific product names (Estee Lauder) are correctly associated with the "Cosmetic" concept.
Experiments and Results
The researchers tested their framework on a dataset of 1,000 certified celebrities from Sina Weibo. They compared their method against typical baselines like Screenname N-grams and raw Tag N-grams.
Performance comparison:
| Feature Method | Accuracy (%) |
|---|---|
| Screenname (Char 1-gram) | 63.67 |
| Tag (Char 1-gram) | 68.33 |
| Tag Vector (Expanded Conceptual Class) | 71.33 |
Table 1: The effect of minimum support () on expansion. Choosing an optimal (5%) ensures the conceptual classes remain relevant without becoming noisy.
Why it Works
Raw tag vectors (without CC) achieved only 65.33% accuracy. By introducing the Original Conceptual Class, accuracy jumped to 70.67%, and the Expanded version pushed it even further. This proves that "densifying" the feature space by grouping individual tags into semantic buckets is the key to handling sparse user data.
Critical Insight & Conclusion
This paper highlights a critical shift in User Profiling: when behavioral data (posts) is missing, identity data (tags) becomes the primary signal.
Key Takeaways:
- Sparsity is the Enemy: Standard N-gram methods fail when users have only 2-3 features.
- Association Mining as a Bridge: Leveraging unlabeled data to expand small "Seed" sets allows models to understand synonymous or related interests.
- Future Path: While effective, the 27 classes were manually defined. Future work could benefit from unsupervised topic modeling (like LDA) or Word2Vec embeddings to discover these conceptual classes automatically.
For advertisers and developers, this study provides a blueprint for understanding the "silent majority" of their user base using nothing but a handful of account tags.
