EIS: Unmasking Spammers by Linking Identities in Voice Networks
Early identification of spammers through identity linking, social network and call features
This paper introduces EIS (Early Identification of Spammers), a novel system designed to detect malicious actors in VoIP and voice networks who frequently change identities to evade detection. The core method utilizes a weighted call graph and identity linking to aggregate global reputation, successfully identifying spammers even when they "whitewash" their accounts.
Executive Summary
TL;DR: Spammers in VoIP networks evade traditional blacklists by constantly cycling through new identities. This paper presents EIS (Early Identification of Spammers), a framework that "connects the dots" by linking multiple identities belonging to the same physical person. By analyzing weighted call graphs—focusing on who people call and how long they talk—the system aggregates a global reputation that unmasks malicious actors even after they switch to fresh, "clean" IDs.
The Academic Landscape: This work bridges the gap between social network identity resolution and telecommunication security. It is a pioneering effort in using weighted structural-context similarity specifically for anti-spam measures in resource-constrained voice environments.
Problem & Motivation: The "Identity Washing" Loop
In modern telephony, acquiring a new number or VoIP ID is trivial and cheap. This enables a devastating tactic:
- A spammer uses an identity until its reputation drops.
- They discard the identity and activate a new one.
- Because the detection system sees a "new" user with no history, the spammer continues their activity.
Why is this hard to solve? In social networks like Facebook or LinkedIn, we can link accounts using profile names, bios, or photos. In VoIP, we only have Call Detail Records (CDRs). Content-based analysis (listening to calls) is a privacy nightmare and computationally impossible at scale. The authors’ insight is that while a spammer can change their ID, they cannot easily change their victim pool or their interaction patterns.
Methodology: The Core of EIS
The EIS system operates through a three-stage pipeline:
1. ID-CONNECT (The Tracker)
This module builds a weighted call graph where nodes are identities and edges represent call interactions. The weight () is derived from the average call duration, serving as a proxy for relationship strength.
Two identities are linked if they share mutual friends and exhibit similar calling behaviors toward those friends.

2. Reputation Aggregation (The Judge)
Once identities are linked, EIS doesn't just look at one ID; it looks at the Physical Individual. It uses an iterative algorithm (similar to PageRank) to compute a global reputation score based on interaction frequency, duration, and out-degree across all linked accounts.
3. Dynamic Thresholding (The Executioner)
Instead of a fixed cutoff, EIS uses a percentile-based approach. It analyzes the 25th percentile of the reputation distribution to distinguish between legitimate users and the "cluster" of spammers who typically occupy the lower end of the trust spectrum.
Experiments & Critical Results
The authors tested EIS against three random graph models: Barabási–Albert (BA), Erdös Réenyi (ER), and Watts–Strogatz (WS).
Higher Precision, Smaller Search Space
Traditional Jaccard similarity (OMF) often returns a large "candidate set" of potential matches, which is noisy. ID-CONNECT reduced the candidate set size by over 50% while maintaining the same True Positive Rate (TPR). By incorporating call duration into the similarity math, it successfully filtered out coincidental mutual friends.

The Battle Against Whitewashing
As shown in the spam detection results, systems without identity linking see a rapid decay in detection rates as spammers rotate IDs. EIS, however, maintains high detection accuracy because it tracks the individual behind the shroud of multiple identities.

Critical Analysis & Conclusion
Takeaway
The true value of this paper lies in its movement away from identity-based security toward behavioral-identity security. It proves that even in metadata-poor environments like telephony, structural call patterns are unique enough to serve as a fingerprint.
Limitations
- Low Overlap Scenarios: If a spammer is sophisticated enough to vary their victim list completely between identities, the TPR drops to around 20%.
- Large Scale Processing: Computing weighted similarity across millions of nodes in real-time remains a significant engineering hurdle not fully addressed in the synthetic evaluation.
Future Outlook
As botnets become more sophisticated, integrating this structural analysis with Speech Fingerprinting (processing only the noise/frequency characteristics of a voice rather than the content) could provide a multi-factor defense that is nearly impossible for spammers to bypass.
