The Digital Echo: Why Your Display Names Are More Traceable Than You Think

Understanding the User Display Names across Social Networks

2017-01-01
Yongjun Li, You Peng, Zhen Zhang, Quanqing Xu, Hongzhi Yin
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comprehensive measurement study on user display names across Foursquare, Facebook, and Twitter, introducing a distributed crawling framework to analyze "redundant information." The study discovers that over 45% of users maintain identical display names across platforms, achieving a state-of-the-art behavioral profile for cross-platform user identification.

TL;DR

Researchers from Northwestern Polytechnical University have quantified the "redundancy" in how we name ourselves online. By analyzing 1.3 million accounts across Facebook, Twitter, and Foursquare, they found that nearly half of us use identical names everywhere, and even those who don't leave behind a "letter distribution" fingerprint that is statistically distinct from other users.

The Motivation: The Memory Trap

Why do we use the same names? The study pivots on a psychological Inductive Bias: Human memory limitations. We tend to use consistent behaviors to make our digital presence easier to remember or to build a unified online reputation.

While previous research heavily focused on usernames (the unique, often alphanumeric handles), this paper argues that display names (the "friendly" names) are actually more valuable. In many networks like Foursquare or QQ, usernames are just random strings of numbers, making display names the primary conduit for human-readable identity.

Methodology: Mining the Cross-Site Link

The authors exploited the "Cross-site linking" function of Foursquare—where users voluntarily link their Facebook and Twitter profiles. This provided a rare "Golden Dataset" where identities are already confirmed, allowing for a rigorous comparison between "Positive" pairs (the same person) and "Negative" pairs (different people).

Figure 1: Two Foursquare user's public profile pages

The Three Dimensions of Similarity

To quantify just how similar these names are, the team applied three specific metrics:

  1. Character Similarity: Using LCS (Longest Common Substring) and Edit Distance. They found that if the edit distance between two names is less than half the total length, there is a very high probability they belong to the same person.
  2. Visual Structure (Best Match): Segmenting names into <First, Middle, Last> and checking if parts are swapped or omitted.
  3. Statistical Fingerprinting: Using Jensen-Shannon (JS) Divergence to compare the probability distribution of the 26 English letters. Even if you change "John Doe" to "Doe John," your JS similarity remains nearly 1.0.

Key Insights from the Data

The study revealed several SOTA observations regarding user behavior:

  • The 45% Rule: Between 45% and 63% of individuals use the exact same display name across platforms.
  • Real-Life Mirroring: The letter distribution in OSN display names almost perfectly mirrors the distribution of common names in real-life census data (peaks at 'a', 'e', 'n', 'i').
  • Platform Specificity: Names on Facebook and Foursquare are more similar to each other than to Twitter, likely because the former two encourage "real-name" usage while Twitter leans toward "handles."

Figure 2: Letter Distribution Comparison with Real Names

Persistence over Time

One of the most significant contributions of this work is the Evolution Analysis. By dividing the data into nine chronological chunks based on registration ID, the authors proved that these naming behaviors are time-independent. Whether you joined a network in 2010 or 2016, the redundancy of your display names remains constant.

Figure 11: Evolution analysis results across chronological datasets

Critical Analysis & Takeaways

This paper provides the mathematical "glue" for future cross-site user identification systems.

Takeaway 1: Privacy is an illusion of effort. Most users who try to change their names for privacy only change parts of them (e.g., omitting a last name), which does little to stop modern string-matching algorithms.

Takeaway 2: Letter distribution is an underrated feature. The fact that JS similarity for positive pairs is consistently above 0.8 while negative pairs stay below 0.7 suggests that "letter frequency" can serve as a lightweight pre-filter for large-scale identity matching.

Limitations: The study focuses primarily on the 26-letter English alphabet. As social networks grow in non-Western regions, the metrics of LCS and Edit Distance may need to be adapted for multibyte character sets or phonetic transcriptions of non-Latin names.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize display name similarity and machine learning classifiers to automate user identity linkage across heterogeneous social networks.
  • Which 2013 paper by Zafarani and Liu first established the "MOBIUS" behavioral-modeling approach for connecting users, and how does this study's display name analysis build upon that foundation?
  • Are there studies that explore how display name redundancy varies across phonetic-based languages (like English) versus logographic languages (like Chinese or Japanese)?
Contents
The Digital Echo: Why Your Display Names Are More Traceable Than You Think
1. TL;DR
2. The Motivation: The Memory Trap
3. Methodology: Mining the Cross-Site Link
3.1. The Three Dimensions of Similarity
4. Key Insights from the Data
5. Persistence over Time
6. Critical Analysis & Takeaways