Cracking the Identity Code: Enhancing Chinese Anchor User Identification via Phonetic and Visual Intuition
Display Name-Based Anchor User Identification across Chinese Social Networks
This paper introduces a display name-based anchor user identification method specifically optimized for Chinese social networks. By extracting unique pronunciation (Pinyin) and font-based (Wubi) features and utilizing the Gradient Boosting algorithm, the authors achieve State-of-the-Art (SOTA) performance in mapping common users across platforms like Sina Weibo, Zhihu, and Douban.
TL;DR
Identifying "anchor users"—individuals who maintain accounts across multiple social networks—is foundational for cross-platform information studies. However, English-centric algorithms often fail on Chinese data. This paper introduces a sophisticated identification model that leverages the unique phonetic (Pinyin) and structural (Wubi) characteristics of Chinese characters, achieving a significant performance boost over standard string-matching baselines using Gradient Boosting.
Background & Motivation: Why Chinese Social Networks are Different
In Western social networks, identifying users via "Screen Names" is relatively straightforward. However, the authors identify two major hurdles in the Chinese digital ecosystem:
- Platform Constraints: Platforms like Sina Weibo require unique display names, forcing users to append random characters or change their preferred handles (e.g., adding "123" or "VIP").
- Linguistic Complexity: Chinese users often swap characters with the same sound (homophones) or similar visual structures (glyphs) when their primary choice is taken or for stylistic reasons.
Existing models based purely on Edit Distance or context similarity ignore these phonological and visual "near-misses", leading to high false-negative rates.
Methodology: Beyond Simple Strings
The core innovation lies in the extraction of four domain-specific features that transform Chinese characters into a machine-interpretable "Identity DNA":
1. The Phonetic Layer (Initial, Final, and Soundex)
By converting Chinese names into Pinyin, the model can calculate:
- Initial/Final Similarity: Capturing rhymes and similar starting sounds.
- Pinyin Soundex: Encoding the pronunciation into a standard phonetic code. This allows the model to recognize that "啊咔呗啦" and "啊咔贝拉" are likely the same person because they sound identical, even if the characters differ.
2. The Structural Layer (Wubi)
To resolve the ambiguity of homophones (different characters that sound the same), the authors used Wubi encoding. Wubi is based on the strokes and visual components of characters. If two names look visually similar but sound different (or sound the same but look different), Wubi provides the necessary tie-breaker.
Fig 1: Illustrating how platform settings (Sina Weibo) force variations that break standard similarity metrics.
Experiments and Results
The authors collected a ground-truth dataset via a custom crawler targeting Sina Weibo, Zhihu, and Douban. They tested 12 machine learning classifiers, identifying Gradient Boosting (GB) as the superior engine for this task.
Performance Gains
The model was tested on two primary datasets:
- D(S,Z) (Weibo-Zhihu): Achieved an F1-score of 0.926, outperforming the previous state-of-the-art by Li Y et al.
- D(S,D) (Weibo-Douban): Achieved an F1-score of 0.752. The lower score here reflects the higher diversity in naming conventions across these specific platforms but still represents a 2.79% improvement over baselines.
Fig 2: Ablation study showing that adding Wubi and Phonetic features consistently improves F1-scores across different datasets.
Key Insight: The Power of Wubi
Interestingly, in the Weibo-Zhihu dataset, Wubi similarity proved to be the most potent individual feature. This suggests that users on these platforms often utilize visual variations of characters when their original names are unavailable.
Critical Analysis & Conclusion
Takeaways
This research proves that "Linguistic Inductive Bias" matters. By baking the rules of Chinese phonetics and writing into the feature engineering process, the model captures human-centric naming patterns that general-purpose algorithms miss.
Limitations & Future Work
The model currently excels at identifying users with similar names. However, "Bridge Users"—those who use completely different identities (e.g., "DeepTech_Expert" on Zhihu vs. "CoffeeLover99" on Douban)—remain elusive. The authors suggest that future work will need to integrate behavioral data (posting times, topics) to link these disparate identities.
Final Verdict
For technologists working on social graph alignment or cybersecurity in Asian markets, this paper provides a vital blueprint: Don't just look at the characters; listen to how they sound and look at how they are built.
