ItemSpider: Turning Your Bookshelf into a Social Compass
ItemSpider: Social Networking Service That Extracts Personal Character from Individual’s Book Information
ItemSpider is a novel Social Networking Service (SNS) that extracts personal characteristics from users' private book collections using Amazon Web Services (AWS). It leverages vector-based user attributes derived from book metadata (authors, publishers, categories) to cluster similar users and foster new community formation.
TL;DR
ItemSpider is a hybrid social network that uses your reading habits as a "personality fingerprint." By extracting metadata from your book collection through AWS, the system creates vector representations of your interests, allowing it to cluster you with like-minded bibliophiles and automate the discovery of new communities.
Academic Positioning: This work bridges the gap between Personal Information Management (PIM) and Social Networking Services (SNS), moving from manual profile filling to automated identity extraction through consumer data.
Problem & Motivation: The "Empty Profile" Trap
Most social networks suffer from a cold-start or static profile problem. Users rarely update their demographics, and "interests" are often limited to shallow tags. The authors recognized that book collections are an honest reflection of an individual's character.
However, existing tools like "Booklog" or "Delicious Library" were silos. They helped you organize your books, but they didn't help you find the people whose libraries mirrored your soul. The challenge lies in turning a list of ISBNs into a mathematically comparable "user attribute" that a machine can use for community building.
Methodology: From ISBNs to Interest Vectors
ItemSpider's architecture is divided into two main engines: Information Gathering and Information Analysis.
1. Feature Extraction via AWS
The system uses the Amazon Web Service (AWS) API to transform a simple ISBN into a rich data packet including Title, Author, and—most importantly—BrowseNode categories. These categories (27 in total, such as "Computer & Internet" or "Literature & Criticism") form the basis of the user's feature vector.
2. The Math of Similarity
To determine if two users should be in the same "community," the system treats each user as a point in a 27-dimensional space.
- Euclidean Distance: Measures the absolute gap between category counts.
- Cosine Similarity: Measures the orientation of interests, ensuring that two users who love the same genres (even if one has more books than the other) are recognized as similar.

Experiments: Testing the "Overlap" Hypothesis
The researchers conducted an experiment on 12 real users and 96 sample users. They wanted to see if the K-means clustering algorithm would naturally group people who shared a certain percentage of the same books (20% to 80% overlap).
Key Findings:
- The 80% Threshold: When users shared 80% of their books, the system placed them in the same cluster 79% of the time.
- The "Same Cluster" Premium: In every scenario, the similarity score for users in the same cluster was consistently higher than those in different clusters, proving the logic of the vector space model.

Critical Analysis: The Limits of Clustering
While the system is effective, the authors noted an interesting technical hurdle: K-means is a "hard" clustering algorithm. Because it forces a user into a single group, two users with nearly identical libraries might occasionally end up in different clusters if they sit near a mathematical boundary.
Furthermore, count-based vectors might miss the nuance of a user's preference. Buying a book isn't the same as liking it. This suggests that future SNS models should incorporate "Soft Clustering" (where you can belong to multiple groups) and sentiment analysis from "Book Reviews" to refine the user vector.
Takeaway & Future Work
ItemSpider demonstrates that our consumer choices (what we buy and read) are powerful enough to automate social networking. The transition from "who you say you are" to "what you consume" marks a shift toward more intelligent, data-driven social structures. Future iterations plan to integrate live usage data from university students to refine these "intellectual community" formations.
Note: This post is based on the paper "ItemSpider: Social Networking Service That Extracts Personal Character from Individual’s Book Information".
