Personalizing Knowledge: Using SVMs to Inject Background Knowledge into Ontologies
Background knowledge for ontology construction
This paper presents an enhancement to the OntoGen system for semi-automatic topic ontology construction by incorporating user-defined background knowledge. The core method introduces a novel word weighting schema derived from SVM feature selection, allowing users to influence how document similarity is calculated based on their specific perspectives.
TL;DR
Ontologies are not universal; a manager and a developer see the same dataset through different lenses. This paper describes an upgrade to OntoGen, a semi-automatic ontology construction tool. By using SVM-based feature selection, the system learns to weight words based on user-provided categories, allowing machine learning algorithms (LSI and k-means) to cluster data in a way that aligns with the user's specific mental model.
The Problem: The Subjectivity Gap in Text Mining
Standard text mining approaches often treat document similarity as a purely statistical property. We typically use TF-IDF (Term Frequency-Inverse Document Frequency) to decide which words are important. While robust, TF-IDF is "blind" to context.
If a user is interested in a geographical breakdown of news, the word "London" should be highly weighted. If they are interested in financial themes, "London" might be less important than "Stock-exchange." Existing systems lacked a mechanism to bridge this gap between statistical significance and user relevance.
Methodology: Turning Classifiers into Weighting Engines
The authors' insight is to use the Support Vector Machine (SVM) not just for classification, but as a feature-importance analyzer.
1. Capturing User Intent
The user provides "Background Knowledge" by labeling a subset of documents. This isn't meant to be an exhaustive manual ontology, but a set of markers that show the system "these documents belong together for my current task."
2. The SVM Weighting Algorithm
Instead of relying on TF-IDF, the system:
- Trains a Linear SVM for each user-defined category.
- Analyzes the Normal Vector () of the SVM. Each dimension in corresponds to a word; a high value indicates the word is a strong predictor for that category.
- Calculates an average "vote" for each word across documents, creating a weighted vector that represents the user's perspective.
3. Concept Discovery
These weights are then injected into the Bag-of-Words representation. When the system performs Latent Semantic Indexing (LSI) or k-means clustering, the "distance" between documents is now warped by the user's background knowledge.
Figure 1: Contrast between clustering based on 'Topic' labels (left) and 'Country' labels (right) for the same dataset.
Experimental Validation
Using the Reuters RCV1 dataset, the authors showed how the same 5,000 documents could yield different ontologies.
- When weighted by Topics, documents about the New York and UK stock exchanges were grouped together under a "Market" concept.
- When weighted by Countries, the same documents were split into "USA" and "UK/Europe" concepts.
This proves that the system successfully "learns" to see the data through the user's specific prism.
Critical Analysis & Future Outlook
While the approach is elegant, it relies on the user providing enough labeled data. The authors suggest Active Learning or using Tagging Services (like the historical Del.icio.us) to mitigate the manual effort.
Takeaway for Today's AI: In the era of RAG (Retrieval-Augmented Generation) and Large Language Models, this paper's philosophy remains highly relevant. While we use embeddings today instead of Bag-of-Words, the challenge of "Personalized Relevance" remains. This work reminds us that the best "Similarity" metric is the one that understands the user's intent.
References
- Fortuna, B., et al. (2006). Background knowledge for ontology construction. WWW '06.
