Towards an Inductive Methodology for Ontology Alignment Through Instance Negotiation
Towards an Inductive Methodology for Ontology Alignment Through Instance Negotiation
The paper introduces an inductive methodology for ontology alignment based on instance negotiation and machine learning. It proposes the "YIN YANG" algorithm to learn concept definitions in Description Logic (ALC) by exchanging shared individuals between autonomous agents.
TL;DR
This paper presents a shift from static, centralized ontology merging to a dynamic, agent-based Instance Negotiation protocol. By treating ontology alignment as a Machine Learning problem, the authors allow software agents to learn what a foreign concept means by looking at shared examples (individuals) and defining them using their own internal "vocabulary" (ontological primitives).
Problem & Motivation: The Failure of Centralization
The original vision of the Semantic Web relied on shared ontologies. However, in the real world, "High Quality" to a supplier means "durable materials," while to a customer it means "comfort and support."
Current SOTA methods like GLUE or PROMPT often require:
- Significant human intervention.
- Global mapping of the entire ontology (which is often overkill).
- Hardcoded linguistic similarities that fail across different languages or naming conventions.
The authors argue that alignment should be on-demand, limited in scope, and automatizable. Instead of forcing everyone to use the same dictionary, why not teach agents to translate on the fly?
Methodology: The YIN YANG Algorithm
The core of this approach is the Concept Learning problem. When Agent A wants to talk about concept C to Agent B, and B doesn't know C, the following happens:
- Instance Exchange: Agent A sends a set of positive and negative examples (URIs of individuals) representing
C. - Local Interpretation: Agent B looks up these individuals in its own Knowledge Base.
- Inductive Learning: Agent B uses the YIN YANG algorithm to search for a definition in its own ontology that covers all positive examples and excludes all negative ones.
The Search Space & Refinement
The algorithm moves through the search space using two operators:
- Generalization (δ): Making a concept broader to cover more positive examples.
- Specialization (ρ): Making a concept more specific to exclude incorrectly covered negative examples (using Counterfactuals—negated descriptions of what went wrong).
Figure 1: The SELA (Self Explaining and Learning Agents) interaction protocol.
Experiments & Results
The authors tested their prototype, SELA, across three scenarios:
- Artificial Translation: Mapping an English academic ontology to an Italian one. The learner achieved perfect equivalent definitions because the underlying structures were identical.
- OAEI Benchmarks: Mapping the French
onto206toonto101. While successful, it revealed a limitation: the learner could not handle cardinality restrictions (e.g., "must have exactly 1 publisher"), leading to overly broad definitions. - Real-World Data (ACM & DBLP): Learning the concept "Article" across two different bibliographic databases.
Performance Breakdown
- Precision: 100% on regular datasets (like DBLP) with as few as 10 positive examples.
- Scalability: Inference/learning time remained efficient, taking approximately 80 seconds for a problem involving 550 total examples.
Figure 2: Example of different viewpoints for "HighQuality" in the food domain.
Critical Analysis & Conclusion
Takeaway
The genius of this approach is its inductive bias. It doesn't care what you call a concept; it cares how you use it. By focusing on shared instances (individuals), it bypasses the "Tower of Babel" problem where different naming conventions block communication.
Limitations
- Expressivity Gap: By sticking to the ALC logic for tractability, the method misses complex constraints like "at least/at most N" (cardinality).
- Example Dependency: The quality of the learned mapping is only as good as the diversity of the examples exchanged. If the teacher provides "biased" samples, the learner will "overfit."
Future Outlook
The next step for this research is to integrate Similarity Measures to evaluate how close a learned definition is to existing local concepts, potentially allowing for automatic "merging" of discovered concepts into the local schema.
