SHAQAR: Revolutionizing Professional Group Extraction with Ranked Semi-Supervised Clustering
Group extraction from professional social network using a new semi-supervised hierarchical clustering
The paper introduces SHAQAR, a novel semi-supervised hierarchical clustering method tailored for group extraction in professional social networks like LinkedIn. It enhances the clustering process by integrating user-defined, quantitative ranked constraints (scores between 0 and 1) to prioritize specific merging operations based on attributes like "Area of Expertise" or "Job Openings."
TL;DR
In the vast sea of LinkedIn data, forming high-quality professional groups manually is impossible, and standard machine learning often misses the mark. This paper introduces SHAQAR, a semi-supervised hierarchical clustering algorithm that uses quantitative ranked constraints. Unlike traditional methods that treat all rules as equal, SHAQAR allows experts to "score" the importance of different professional criteria, leading to a much more accurate representation of real-world business communities.
Background: Why Professional Networks Change the Game
Professional social networks like LinkedIn differ from general social media (like Facebook) because relationships are defined by specific utility—expertise, recruitment, and industry alignment. Standard clustering algorithms often fail here because they treat every data point’s similarity equally. The authors argue that in a professional context, some attributes (like "Area of Expertise") are inherently more "discriminating" than others (like "Date of Group Creation").
The Core Problem: The Binary Constraint Limitation
Existing semi-supervised clustering usually relies on Boolean constraints:
- Must-Link (ML): Two users must be in the same group.
- Cannot-Link (CL): Two users must not be in the same group.
The Problem? Reality is rarely binary. An expert might strongly prefer merging two software engineers while only moderately preferring to merge two people from the same city. Prior methods like Klein et al. or COP-K-means cannot handle these nuances of priority, often resulting in suboptimal clusters when multiple conflicting constraints exist.
Methodology: The SHAQAR Insight
The SHAQAR (Semi-supervised Hierarchical Active clustering based on QuAntitative Ranking) method introduces three major innovations:
- Numeric Scoring: Instead of True/False, constraints are assigned a score .
- Highest Operator: In each iteration of the hierarchical sequence, the algorithm identifies the constraint with the highest priority to execute the merge.
- Multidimensional Social Warehousing: The data is pre-processed into a social warehouse (OLAP-ready), allowing the algorithm to pull from hierarchies like Expertise, Security, and Job Openings.
Fig 1: The process of building the professional network warehouse used for SHAQAR.
The algorithm uses the Jaccard similarity coefficient to calculate the initial distance between user profiles based on 24-31 distinct characteristics (diplomas, skills, etc.) and then modifies the merging sequence based on the ranked constraints.
Experimental Results: Proving the Advantage
The authors tested SHAQAR against a dataset of 300 LinkedIn users and compared it to established baselines, including the ID3 decision tree classifier and the popular Klein et al. semi-supervised algorithm.
1. Superior Accuracy
Across all four tested criteria (Expertise, Security, Jobs, Time), SHAQAR consistently outperformed previous methods.
Fig 2: Accuracy of SHAQAR vs. Klein et al. based on the "Area of Expertise" criterion.
2. Key Findings
- Expertise is King: The "Area of Expertise" proved to be the most effective criterion for group extraction, yielding a precision of 90% for technical domains.
- Data-Driven Success: Open groups and groups offering "Job Openings" were significantly easier to cluster correctly, as they showed stronger community signals than closed or non-hiring groups.
Deep Insight & Conclusion
The true value of this paper lies in its movement away from "rigid" semi-supervised learning. By allowing experts to rank constraints, the researchers have effectively created a human-in-the-loop system that balances automated data patterns with human intuition.
Limitations: The study is currently conducted on a relatively small sample (300 users). Scaling this "Numeric Scoring" system to millions of users would require a more automated way of generating these scores, perhaps through secondary machine learning models rather than manual expert input.
Future Outlook: Transitioning this from static warehouse data to RSS feeds and streaming data (as the authors suggest) will be the next major hurdle for professional social network analysis.
