DBUL: Reconsidering User Identity Linkage through the Lens of Density-Based Clustering
DBUL: A User Identity Linkage Method across Social Networks Based on Spatiotemporal Data
This paper introduces DBUL, a novel User Identity Linkage (UIL) method using DBSCAN clustering to match users across social networks via spatiotemporal data (FS-TW and IG-TW). By representing scattered check-in points as distinct cluster centers, the method achieves SOTA precision and significantly improves computational efficiency compared to grid-based or trajectory-based baselines.
TL;DR
Connecting the dots between a Foursquare check-in and a Twitter status update isn't just a matter of drawing lines. DBUL (DBSCAN-based User Linkage) revolutionizes this by discarding the rigid "trajectory" and "grid" models. Instead, it uses density-based clustering to extract the "essence" of a user's location habits, delivering SOTA precision and reducing computational overhead by up to 46.5%.
The "Sparsity" Trap in Spatiotemporal Data
User Identity Linkage (UIL) is the backbone of cross-domain recommendation and personalized advertising. However, social network spatiotemporal data is notoriously messy:
- Sparsity: Unlike a continuous vehicle GPS trace, social check-ins are snapshots separated by hours or even days.
- Grid Distortions: Previous SOTA methods like GKR-KDE chop the world into grids. If a user check-ins on the very edge of two adjacent grids, the system fails to recognize them as the same location—the "edge exception" problem.
- Noise: Random one-time check-ins often skew trajectory-matching algorithms.
The authors of DBUL propose a simple yet powerful physical intuition: Humans are creatures of habit. Regardless of which app they use, they gravitate toward a few "cluster centers" (home, office, favorite cafe).
Methodology: From Points to Cluster Centers
DBUL replaces raw data points with a set of cluster centers .
1. The DBSCAN Advantage
Unlike K-Means, DBSCAN doesn't require knowing the number of clusters in advance and can handle irregular shapes. This is perfect for capturing the unique spatial "footprint" of a human user.
2. The Similarity Architecture
The method employs a modified Ochiai Coefficient to calculate similarity. Because two GPS coordinates rarely match perfectly, it introduces a coincidence_radius. If two cluster centers from different platforms are within this radius, they are counted as an intersection.

3. Smart Noise Handling
The paper introduces "Reserve Conditionally" logic. If a user's data is overwhelmingly noisy, the noise points are retained as individual clusters to prevent total information loss. If clear clusters exist, the noise is discarded to maintain efficiency.
Experimental Battleground: FS-TW vs. IG-TW
The researchers tested DBUL against industrial-strength baselines including BIN, DG, and GKR-KDE.
Performance Highlights:
- Precision Supremacy: On the Foursquare-Twitter (FS-TW) dataset, DBUL hit the highest precision among all methods.
- Density Robustness: On the larger Instagram-Twitter (IG-TW) dataset, which is denser and more complex, DBUL swept all metrics (Precision, Recall, F1).

Efficiency Gains:
By reducing thousands of records into a handful of cluster centers, the computational volume drops drastically. DBUL clocked an average running time significantly lower than BIN and GS, making it a viable candidate for real-time large-scale deployments.
Critical Insight & Conclusion
The true value of DBUL lies in its Inductive Bias. By assuming that user behavior is center-driven rather than pathway-driven, it finds a "signal" in the "noise" of sparse social data.
Takeaway: For practitioners dealing with sparse behavioral data, the lesson is clear: don't try to reconstruct the path; find the hubs. While DBUL currently focuses on spatial data, its clustering logic could easily be extended to temporal or even semantic behavioral clusters in future iterations.
