Time is Location: Decoding Twitter Geography via Account Creation Cycles
Improving Geocoding of a Twitter User Group using their Account Creation Times and Languages
The paper introduces a two-step framework for geocoding Twitter influencers and their followers by utilizing universal, language-independent features: account creation times and language preferences. It achieves high-precision country-level location prediction (up to 97.54%) by modeling a group's "sleep cycle" from creation time distributions and refining candidates via language similarity.
TL;DR
Researchers have developed a method to locate Twitter influencers and their communities worldwide without reading a single tweet. By analyzing the account creation times—a feature available for nearly 100% of users—they can identify "sleep cycles" that reveal a group's Time Zone. Coupled with language preferences, this approach boosts geocoding precision to over 97%, even for non-English speaking regions.
Problem & Motivation: The Geocoding Blind Spot
Understanding the geographic reach of an influencer is critical for marketing, news dissemination, and identifying disinformation campaigns. However, geocoding Twitter users is notoriously difficult because:
- Data Scarcity: Only a fraction of users provide valid self-reported locations.
- Linguistic Bias: Most geocoders are optimized for English and the Latin alphabet.
- Privacy: Twitter restricted access to explicit "Time Zone" and "UTC Offset" fields in 2018.
The authors' insight is simple yet profound: No matter what language you speak, you probably aren't creating social media accounts while you sleep. This biological constraint creates a "temporal signature" in the data.
Methodology: The Sleep-Time Signature
The core of the method is the Normalized Time Distribution. By plotting when a group of followers created their accounts, a "U-shaped" dip appears, representing the hours when that population is typically asleep.
1. Identifying the "Peak Sleep Time" (PST)
The researchers use a moving average to smooth the data and then fit a parabola to the lowest 33.3% of the distribution.
- Local vs. Global: If a group shows a clear U-shape, they are likely a "geo-influencer" group concentrated in one region.
- Universal Correlation: As shown in the study, PST has a near-perfect linear relationship () with the actual UTC offset.
Fig 1: Smoothed account creation distributions for Chicago (Top) and London (Bottom) followers.
2. Language Constraint
Since multiple countries share the same time zone (e.g., UK and Nigeria), the system compares the distribution of the 76 supported Twitter languages within the group against known country profiles using Cosine Similarity.
Scaling with an Improved Geocoder
The researchers used their high-confidence predictions to train a multilingual TF-IDF geocoding model. This allowed the system to "learn" how users in specific countries refer to their locations in native scripts (e.g., 日本 for Japan or Москва for Russia) without needing a pre-defined dictionary.
Fig 2: The full pipeline from raw follower data to country-level localization.
Experiments & Results
The framework was tested on a massive dataset of 320,000 influencers and 377 million user profiles.
- Precision Leap: For foreign (non-US) influencers, where traditional geocoders often struggle, the precision increased from 65.71% to 86.59%.
- Point Accuracy: When requiring the highest confidence (S3_Point), the geocoder reached 97.54% precision.
Table 1: Performance comparison showcasing the iterative improvement of the TF-IDF model.
Critical Analysis & Future Outlook
Takeaway
This work demonstrates that "system metadata" (account creation time) is often more reliable than "user-provided content." It provides a privacy-preserving, language-agnostic way to map the digital world.
Limitations
- Resolution: The method currently works best at the country or region level, not at the city or street level.
- Global Influencers: "World" celebrities whose followers are spread across all time zones won't show the characteristic "sleep cycle" dip, making them harder to locate.
Future Work
The authors suggest that by analyzing why certain influencers lack a sleep cycle, we can better rank "Global vs. Local" reach—a vital metric for understanding the true "geographical weight" of a digital voice.
