LinkedIn Skills: Mapping the Global Professional Genome via Folksonomy and Inference

LinkedIn skills: large-scale topic extraction and inference

2014-10-01
Mathieu Bastian, Matthew Hayes, William Vaughan, Sam Shah, Peter Skomoroch, Hyungjin Kim, Sal Uryasev, Christopher Lloyd, Christopher Lloyd
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents LinkedIn's large-scale "Skills and Expertise" system, detailing the creation of a massive professional folksonomy and a recommendation engine. By combining NLP, clustering, and Naive Bayes inference, the system automates skill tagging for hundreds of millions of professional profiles.

TL;DR

LinkedIn revolutionized how professional expertise is captured by moving from a static, expert-curated taxonomy to a data-driven folksonomy. By extracting entities from millions of profiles, disambiguating them via clustering, and leveraging a Naive Bayes recommender system, they boosted skill-tagging conversion by over 1,200%.

Background & Motivation: The Problem of "The Blank Box"

How do you categorize the skills of 300M+ professionals ranging from ballet dancers to nuclear engineers? LinkedIn's initial attempts were thwarted by two main issues:

  1. Taxonomy Scalability: No existing list was comprehensive enough to cover the "long-tail" of specialized professional niches.
  2. The Interaction Gap: Simply providing a search box (Type-ahead) led to stagnant profiles. Users are much more likely to confirm a suggestion than to recall and type one from scratch.

Authors identified that professional identity is often "hidden in plain sight" within free-text profile sections. The challenge was converting this messy text into standardized, deduplicated, and disambiguated "entities."

Methodology: The Three Pillars of Folksonomy

1. Discovery & Entity Extraction

The team focused on the "Specialties" section of profiles, identifying comma-separated lists using a specific punctuation frequency threshold (). This allowed them to filter prose from keywords, yielding 150,000 candidate phrases.

2. Disambiguation: The "Organ" Problem

A skill like "Organ" means something different to a Cardiologist than to a Jazz Musician. To solve this, the authors used Co-occurrence Clustering:

  • Jaccard Similarity: Measured how often two phrases (e.g., "Organ" and "Surgery") appeared together.
  • SVD & KMeans: Applied dimensionality reduction to the similarity matrix to identify distinct "senses" of a word.
  • Industry Labeling: Attached the most common member industry to each cluster to provide context.

Model Architecture Placeholder Figure 1: The Skills & Expertise section UI, the final output of the pipeline.

3. Deduplication via Crowdsourcing

Members use synonyms like "Java Programming" and "Java Development." The authors cleverly mapped these to Wikipedia entities using Amazon Mechanical Turk. If two different phrases mapped to the same Wikipedia URL, they were merged into a single standardized skill.

Inference: Solving the Cold Start

To suggest skills for users who hadn't listed any, LinkedIn built a Naive Bayes Classifier. Instead of looking at past skills, it looks at "Collaborative Attributes": If your peers at "Google" with the title "Software Engineer" all have the skill "Distributed Systems," the system infers that you likely do too.

Experimental Results: Performance and UX

The most striking result wasn't just accuracy, but User Behavior shift.

Experiment Results Figure 2: Type-ahead (4% conversion) vs. Recommendations (49% conversion).

  • Conversion: By switching to a 10-item recommendation list, conversion jumped from 4% to 49%.
  • AUC Performance: The inference engine achieved an AUC of 0.77. Technical skills (e.g., "Hadoop") performed better because they have "tighter" co-occurrence patterns compared to "soft skills" like "Teamwork."

Critical Insight & Conclusion

The genius of the LinkedIn Skills system lies in its recognition that Identity is Social. By using "Endorsements" as a social gesture and "Inference" as a psychological nudge, they turned a data-entry chore into a core part of professional reputation.

Limitations: The system relies heavily on professional homophily (the idea that people with similar titles have similar skills). This can create a "filter bubble" where rare or cross-disciplinary skills are harder to infer.

Future Work: Transitioning from Naive Bayes to more complex Machine Learned models using real-world user feedback (accept/reject signals) will likely close the gap on "soft skill" inference accuracy.

Find Similar Papers

Try Our Examples

  • Find recent papers that improve upon Naive Bayes for attribute inference in social networks using Graph Neural Networks (GNNs).
  • What are the latest SOTA methods for automated taxonomy construction and disambiguation in professional domains beyond the Jaccard/SVD approach?
  • Research studies investigating the "cold-start" problem in tag recommendation systems for professional identity platforms.
Contents
LinkedIn Skills: Mapping the Global Professional Genome via Folksonomy and Inference
1. TL;DR
2. Background & Motivation: The Problem of "The Blank Box"
3. Methodology: The Three Pillars of Folksonomy
3.1. 1. Discovery & Entity Extraction
3.2. 2. Disambiguation: The "Organ" Problem
3.3. 3. Deduplication via Crowdsourcing
4. Inference: Solving the Cold Start
5. Experimental Results: Performance and UX
6. Critical Insight & Conclusion