Beyond Keyword Matching: Finding Experts via Social Network Propagation
Finding Experts Using Social Network Analysis 1
This paper introduces an Expertise Propagation algorithm that leverages Social Network Analysis (SNA) to enhance expert finding in organizational intranets. By utilizing email communication patterns and web page co-occurrences, the method re-ranks candidate experts, achieving significant performance gains on TREC enterprise benchmarks.
TL;DR
In a corporate environment, who you know is often as important as what you know. This paper proposes a method to find experts not just by what they write, but by who they talk to. By building a social graph from emails and web pages, the authors introduce an Expertise Propagation algorithm that allows "expertness" to flow through a network, significantly improving the accuracy of expert discovery tools.
1. The Context: The Social Nature of Expertise
Most search engines treat people like documents. If you search for "Java Programming," the system looks for people who have "Java" written all over their profiles. However, true expertise is social. Real experts are nodes in a network—they lead discussions, participate in email threads, and are co-authored on technical reports.
The authors identified a gap: existing Stage 1 models (document-based) were good at finding potential candidates but lacked the organizational "intuition" to refine that list. Their insight? If Person A is a confirmed expert, their frequent collaborators are likely experts too.
2. Methodology: Building the Social Graph
The core of the paper lies in how associations are quantified. The authors don't just check if two people exist; they measure the strength of their bond through two channels:
A. The Web Page Network (Co-occurrence)
Instead of a simple "yes/no" for candidates appearing on the same page, the authors use Distance-Weighted Co-occurrence (DW CO).
- Intuition: If two names are in the same sentence, they are likely working together. If they are five paragraphs apart, they might just be in the same company newsletter.
- Formula: The strength is the reciprocal of the number of words between name occurrences.
B. The Email Network (Communication Patterns)
Email is the "bloodstream" of an organization. The authors move beyond single emails to Email Threads.
- Why Threads? A single email might just be a CC to an admin. A thread represents a sustained technical discussion. By smoothing single message data with thread data (Equation 7), they solve the "sparse data" problem where many experts don't interact in every single message but remain part of the same core conversation.
The propagation formula: Candidate receives a fraction of the expertise of , weighted by the strength of their association .
3. Expertise Propagation: A "PageRank" for People
The algorithm works in two stages:
- Seed Selection: Pick the Top-N candidates from a standard document-based search.
- Propagation: These seeds act as "expertise sources," pushing their probability scores to their neighbors in the social graph.
Unlike PageRank, which is iterative and global, this approach is Seed-centric. It assumes the initial search gave us a "hint" of the truth, and social analysis is the "refining fire."
4. Key Results & Evidence
The researchers tested this on the TREC Enterprise track, the gold standard for this field.
| Method | TREC 2006 (MAP) | Improvement |
|---|---|---|
| Baseline (Document only) | 0.4592 | - |
| Web Co-occurrence (DW CO) | 0.4700 | +2.3% |
| Email Threads (EC+TC) | 0.5075 | +10.5% |
Critical Finding: The Role of the Seed
The performance is highly sensitive to the "Seed Size" (N). If N is too small, you ignore potential experts; if N is too large, you inject "noise" into the propagation. The optimal range was found to be between 10 and 30 candidates.
Figure 1: Shows how Mean Average Precision (MAP) peaks as we find the ideal balance of seed experts.
5. Critical Analysis & Future Outlook
While the method is powerful, the authors honestly note its lack of robustness to noise. If a non-expert is accidentally included in the seed, they will "pollute" the network by elevating their equally non-expert friends.
Furthermore, Query-Dependent Social Networks (building a graph only from topic-relevant emails) actually performed worse than global ones in some tests. This is a classic "Small Sample Problem"—by filtering for topics, the data becomes too sparse to form a reliable graph.
Conclusion
This work is a precursor to modern knowledge graphs used in tools like LinkedIn or Microsoft Viva. It proves that in the quest to "know who knows what," the network is often louder than the individual. Future work in this space will likely involve Deep Graph Learning to better filter the noise that this early propagation model struggled with.
