Leveraging the Social Graph: How LinkedIn Resolves Company Entities at Scale
Entity Resolution Using Social Graphs for Business Applications
This paper introduces a robust Entity Resolution (ER) framework for LinkedIn to map member-entered company names to canonical business entities. It utilizes a binary classification approach leveraging social graph features, behavior signals, and content heuristics, successfully scaling to hundreds of millions of records using Hadoop.
Executive Summary
TL;DR: LinkedIn researchers developed a machine learning framework that moves beyond simple text matching to resolve ambiguous company names (e.g., "Orion" could be one of five different firms). By treating a member's professional network as a feature set—relying on the principle that you are likely connected to your real colleagues—they achieved 97% precision in mapping "messy" user data to canonical entities.
Background Positioning: This work represents a critical shift from traditional database deduplication to Social-Aware Entity Resolution. It sits at the intersection of Information Extraction and Social Network Analysis (SNA), providing the backbone for LinkedIn's ad targeting and recruitment products.
The Problem: The "IBM" vs. "IBM Almaden" Dilemma
Entity Resolution (ER) in a semi-structured environment faces two primary hurdles:
- k-Ambiguity: A common name like "CHI" might refer to the "Catholic Health Institute" or "CHI X Networks."
- k-Variance: A single company like "IBM" has over 6,000 recorded variations in member profiles, ranging from official names to specific research centers or acquired subsidiaries.
Standard "Typeahead Assist" tools help, but users often bypass them due to latency or because their specific sub-organization isn't listed. This results in a massive "long tail" of unresolved positions that hurts revenue-critical features like recruiter searches and ad targeting.
Methodology: Beyond String Similarity
The authors propose a logic that hinges on Homophily—the sociological observation that "birds of a feather flock together." If your connections work at "Google," and you type "G-Search," there is systemic evidence that you are likely at Google.
The 3-Step Pipeline
- Candidate Set Construction: Instead of comparing a position to millions of companies, the system limits the search to:
- Companies where the user's first-degree connections work.
- Companies with similar "stems" (high IDF keywords).
- Classification & Selection: A Binary Logistic Regression model evaluates the (Position, Company) pair. Unlike ranking models, this approach allows a position to remain "unresolved" if evidence is insufficient, preventing false positives.
- Sanity Check: Automatic triggers flag anomalies, such as a company's size doubling overnight due to a resolution error.
Fig 1. Identifying organizational clusters through network visualization (InMaps).
The Feature Set
The classifier uses three unique dimensions:
- Content/Demographic: Name rarity (IDF), location matches, and email domain overlaps.
- Social Graph: The raw number of connections a member has at the candidate company.
- Social Behavior: The number of connection invitations received from members at that company—a signal that often proves more accurate than static connection counts for active networkers.
Experiments and Results
The model was tested against a baseline using only name features. While the baseline performs well on "easy" cases (where users selected from a dropdown), it collapses on the "unresolved" long tail.
Fig 2. Precision-Coverage curve showing the Social Graph approach (solid line) vastly outperforming the name-only baseline.
Key Metrics:
- Precision: Reached 97% on manually labeled unresolved data.
- Coverage: Effectively resolved 50% of previously "dark" positions.
- Efficiency: Capable of processing 100 million profiles within hours using Hadoop.
Critical Insight & Conclusion
The true value of this research lies in its exposure of Wisdom of the Crowd. By aggregating metadata from members who did correctly resolve their positions, LinkedIn can "backfill" missing data (like location and size) for the entire company entity.
Takeaway: In modern AI systems, the "Social Context" is often just as informative as the "Content" itself. For business applications, leveraging your user graph to clean your data is not just an optimization—it is a requirement for SOTA performance.
Future Outlook: The authors suggest that this same mechanism could be used for Company De-duplication, effectively letting the social network "merge" duplicate company pages if they share the same employee clusters.
