Disambiguating Search: Leveraging Social Streams for Intent Discovery
Disambiguating Search by Leveraging a Social Context Based on the Stream of User’s Activity
This paper introduces a social-context-driven query expansion method to disambiguate short web searches. By constructing an evolving social network from users' activity streams and detecting overlapping communities, the system implicitly infers session-based context to enrich ambiguous queries with relevant keywords.
TL;DR
This research tackles the "short query" problem—the tendency for users to provide vague search terms like "CSS" or "Java"—by looking at the people around them. By monitoring web activity through a proxy, the system builds a dynamic social network and uses the behavior of similar "communities" to automatically add context-heavy keywords to a user’s search.
Background: The Short Query Paradox
Web search is a battle against ambiguity. A programmer searching for "cucumber" is looking for a testing framework, while a chef wants a salad ingredient. Standard search engines rely heavily on global popularity (PageRank), which often ignores the individual's professional or social context. The authors argue that since users are unlikely to change their habit of writing short queries, the system must do the "heavy lifting" of disambiguation.
Problem & Motivation: Beyond Static Profiles
Previous attempts at personalization often used static user profiles or pre-defined categories (like ODP). However, human interest is fluid. The key insight of this paper is that context is social. If we can identify which "virtual community" a user currently belongs to—based on their recent browsing stream—we can use the collective behavior of that community to predict the true intent behind a vague keyword.
Methodology: From Proxy Logs to Community Intelligence
The system follows a sophisticated pipeline to transform raw web traffic into search context:
- Activity Tracking: Using a proxy server, the system monitors "implicit ratings"—calculated by comparing the time spent on a page vs. the page size.
- Social Graph Construction: A network is built where edges represent similarity. High weights are assigned if two users visit the same specific documents or share significant content features (keywords/tags).
- Community Detection: An activation-energy algorithm identifies overlapping communities. This is crucial because a single user can be both a "programmer" and a "cat lover."
- Query Expansion: When a search is detected, the system pulls keywords from the community’s "Query Streams" (history of refined searches) or "Keyword Co-occurrences."

Experiments & Results: Making "Jaguar" Specific
The authors validated their approach using the massive AOL search dataset and their own proxy logs. The results showed a clear ability to distinguish between different meanings of the same word based on community context.
| Original Query | Expanded (Contextual) Query | Intent Captured |
|---|---|---|
| jaguar | jaguar animal | Zoology / Nature |
| branch | branch git | Software Development |
| apache | apache server | Web Infrastructure |
By analyzing "Query Streams"—the sequence of queries a user makes before finding a result—the system learns that a user who starts with "java" and finishes with "history of indonesia" belongs to a different context than one who finishes with "java runtime environment."

Critical Analysis & Conclusion
The beauty of this method lies in its unobtrusiveness. The user doesn't have to fill out a profile or manually tag their interests; the "Social Context" is built silently in the background.
Limitations:
- Cold Start: New users or niche topics without established community data may not see immediate benefits.
- Privacy: Relying on a proxy to track all user activity raises significant privacy concerns that would need modern encryption/anonymization (like Differential Privacy) to be viable today.
Future Outlook: The transition from manual keyword extraction to Tagsonomies (hierarchical keyword relations) suggests a path toward more semantic understanding. In the modern era, this logic could be integrated into LLM-based search agents to provide even more granular personalization.
