Scaling the Semantic Web: Automatic Integration via Domain Taxonomy Ontology
Automatic Searching Method Based on Domain Taxonomy Ontology
This paper proposes an automatic searching method based on Domain Taxonomy Ontology to integrate disparate information sources on the web. By utilizing a "widest taxonomy" mediator and mathematical operations (A and B operations) for attribute mapping, the system provides users with Exact, Estimative, and Recommendatory answers across heterogeneous databases.
TL;DR
The explosion of web-based databases has made manual information integration a bottleneck. This paper introduces an automated framework that replaces manual mapping with a Domain Taxonomy approach. By representing data as attribute-value vectors, the system automates source-to-mediator alignment and introduces a multi-tier answering model providing Exact, Estimative, and Recommendatory results.
Background: The Limits of Manual Mapping
In the early days of the Semantic Web, "Mediators" acted as the middle layer between users and databases. However, connecting a new database required researchers to manually define subsumption relationships (e.g., Is "Benz" a "Car"?). In a web with millions of sources, this manual labor is the primary obstacle to a unified knowledge search.
The authors identify that finding relationships between conceptions is hard, but finding relationships between attributes (like Brand, Model, Year) is much simpler and can be computed automatically.
Methodology: Taxonomy as a Mathematical Vector
The core innovation lies in the transition from abstract concepts to a Widest Taxonomy.
1. The Domain Taxonomy
The system first defines a global domain (e.g., Automobiles) as a collection of attributes . Each value under these attributes is assigned an integer. A specific car then becomes a vector, such as (1, 2, 0), where 1 might represent "Benz" and 2 represents "Sports Car."
2. System Architecture
The architecture consists of a global taxonomy ontology and multiple local information sources. The taxonomy ontology acts as the "Standard Interface" that unifies access.

3. Automated Mapping through A & B Operations
To automate the search, the paper defines two mathematical operations:
- Operation A (Precise): Finds exact matches where the source attributes perfectly align with the query.
- Operation B (Estimative): Handles "vague" queries where some attributes are unknown (marked as -1). It calculates the "Whole Set" by finding terms in the taxonomy that provide values for those unknown gaps while maintaining similarity in known attributes.

Experiments & Results: Beyond "Yes or No" Answers
The researchers validated their method using an automobile domain instance. When a user queries for "Volkswagen Sedan" but doesn't specify the "Box" type:
- Exact Answer (): Returns only the sources that explicitly map to the "Volkswagen Sedan" node.
- Estimative Answer (): Traverses the tree to find related nodes (like 2-door or 4-door variants) that the user might find relevant.
- Recommendatory Answer: This is a meta-search feature. It calculates a "Repeated Rate"—if multiple indexed sources contain a specific term (e.g., a popular model found in five different car sales databases), it is prioritized as a high-reliability recommendation.
| Type | Set | Result (Indices) |
|---|---|---|
| Exact | Precise(t) | {1} |
| Estimative | Whole(t) | {1, 2, 4} |
Critical Insight: The Value of Categorical Logic
The significance of this work is its Inductive Bias: by assuming that most data can be categorized into a fixed set of attributes, it moves the problem of "Information Retrieval" into the realm of "Set Theory."
Limitations & Future Work
While the automatic mapping is elegant, it still requires a Domain Taxonomy to be defined upfront. The "Widest Taxonomy" can also lead to a "Sparsity Problem" where many attribute combinations in the mediator don't actually exist in reality. Future research could look into using LLMs (Large Language Models) to generate these taxonomies dynamically, further reducing the need for initial domain engineering.
Conclusion
By shifting the focus from manual logic-building to automated taxonomic annotation, this paper provides a robust blueprint for unifying the "Deep Web." The introduction of the "Recommendatory Answer" adds a layer of trust and cross-validation that is often missing in standard keyword-based search engines.
