Personalizing Knowledge: Using SVMs to Inject Background Knowledge into Ontologies

Background knowledge for ontology construction

2006-05-23
Blaz Fortuna, Marko Grobelnik, Dunja Mladenic
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an enhancement to the OntoGen system for semi-automatic topic ontology construction by incorporating user-defined background knowledge. The core method introduces a novel word weighting schema derived from SVM feature selection, allowing users to influence how document similarity is calculated based on their specific perspectives.

TL;DR

Ontologies are not universal; a manager and a developer see the same dataset through different lenses. This paper describes an upgrade to OntoGen, a semi-automatic ontology construction tool. By using SVM-based feature selection, the system learns to weight words based on user-provided categories, allowing machine learning algorithms (LSI and k-means) to cluster data in a way that aligns with the user's specific mental model.

The Problem: The Subjectivity Gap in Text Mining

Standard text mining approaches often treat document similarity as a purely statistical property. We typically use TF-IDF (Term Frequency-Inverse Document Frequency) to decide which words are important. While robust, TF-IDF is "blind" to context.

If a user is interested in a geographical breakdown of news, the word "London" should be highly weighted. If they are interested in financial themes, "London" might be less important than "Stock-exchange." Existing systems lacked a mechanism to bridge this gap between statistical significance and user relevance.

Methodology: Turning Classifiers into Weighting Engines

The authors' insight is to use the Support Vector Machine (SVM) not just for classification, but as a feature-importance analyzer.

1. Capturing User Intent

The user provides "Background Knowledge" by labeling a subset of documents. This isn't meant to be an exhaustive manual ontology, but a set of markers that show the system "these documents belong together for my current task."

2. The SVM Weighting Algorithm

Instead of relying on TF-IDF, the system:

  • Trains a Linear SVM for each user-defined category.
  • Analyzes the Normal Vector () of the SVM. Each dimension in corresponds to a word; a high value indicates the word is a strong predictor for that category.
  • Calculates an average "vote" for each word across documents, creating a weighted vector that represents the user's perspective.

3. Concept Discovery

These weights are then injected into the Bag-of-Words representation. When the system performs Latent Semantic Indexing (LSI) or k-means clustering, the "distance" between documents is now warped by the user's background knowledge.

Conceptual Result Comparison Figure 1: Contrast between clustering based on 'Topic' labels (left) and 'Country' labels (right) for the same dataset.

Experimental Validation

Using the Reuters RCV1 dataset, the authors showed how the same 5,000 documents could yield different ontologies.

  • When weighted by Topics, documents about the New York and UK stock exchanges were grouped together under a "Market" concept.
  • When weighted by Countries, the same documents were split into "USA" and "UK/Europe" concepts.

This proves that the system successfully "learns" to see the data through the user's specific prism.

Critical Analysis & Future Outlook

While the approach is elegant, it relies on the user providing enough labeled data. The authors suggest Active Learning or using Tagging Services (like the historical Del.icio.us) to mitigate the manual effort.

Takeaway for Today's AI: In the era of RAG (Retrieval-Augmented Generation) and Large Language Models, this paper's philosophy remains highly relevant. While we use embeddings today instead of Bag-of-Words, the challenge of "Personalized Relevance" remains. This work reminds us that the best "Similarity" metric is the one that understands the user's intent.

References

  • Fortuna, B., et al. (2006). Background knowledge for ontology construction. WWW '06.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend the OntoGen system or use active learning for semi-automatic ontology construction in the 2020s.
  • Which research first introduced the use of SVM normal vectors for feature selection in text mining, and how does this paper's weighting formula differ from the original?
  • Explore how background knowledge-driven word weighting is being applied to modern Large Language Model (LLM) embedding spaces for personalized document retrieval.
Contents
Personalizing Knowledge: Using SVMs to Inject Background Knowledge into Ontologies
1. TL;DR
2. The Problem: The Subjectivity Gap in Text Mining
3. Methodology: Turning Classifiers into Weighting Engines
3.1. 1. Capturing User Intent
3.2. 2. The SVM Weighting Algorithm
3.3. 3. Concept Discovery
4. Experimental Validation
5. Critical Analysis & Future Outlook
5.1. References