From Art to Craft: Professionalizing Ontology Engineering in the Data Mining Era
Ontology Engineering: From an Art to a Craft - The Case of the Data Mining Ontologies
The paper advocates for the transformation of ontology engineering from an "art" to a methodical "craft." It introduces the OntoDM suite—a modular set of ontologies including OntoDM-core, OntoDT, and OntoDM-KDD—designed to standardize the representation of data mining entities, datatypes, and the knowledge discovery process.
TL;DR
Despite two decades of progress, building an ontology is often seen as a "craftsman's intuition" rather than a standardized engineering discipline. This paper argues for a paradigm shift, utilizing the authors' experience with OntoDM (Ontologies for Data Mining) to highlight the critical need for rigorous testing, objective evaluation, and centralized repositories in the IT domain.
The "Artistic" Bottleneck of Ontology Development
In fields like Machine Learning, we have clear benchmarks (like Accuracy or F1-score) to compare algorithms. In Software Engineering, we have "Unit Testing" and Git for version control. However, Ontology Engineering remains stuck in a subjective phase where:
- Evaluation is mostly based on "Expert Opinion."
- Testing for logical entailment errors is rarely reported.
- Reuse is hampered by a lack of a central hub (an "IT Portal") for non-biological ontologies.
The authors argue that without moving toward a true engineering mindset, ontologies will remain isolated "controlled vocabularies" rather than the powerful "building blocks" of intelligent systems.
Methodology: The OntoDM Ecosystem
To address these gaps, the authors developed a modular ontological suite for the data mining domain. By following the Basic Formal Ontology (BFO) as a top-level template and adopting the MIREOT principle for referencing external terms, they created three core components:
- OntoDM-core: Represents specification, implementation, and application layers of data mining.
- OntoDT: A specialized ontology for datatypes based on ISO standards.
- OntoDM-KDD: Represents the Knowledge Discovery process based on the CRISP-DM model.
(Note: This illustrates the transition from ad-hoc design to the proposed systematic evaluation and testing framework.)
The Core Challenge: Why Reasoning Isn't Enough
A recurring theme in the paper is the distinction between Evaluation (does it meet requirements?) and Testing (finding errors). A major insight is that standard Reasoners (like HermiT or Pellet) only check for consistency.
For instance, a Reasoner won't complain if a "Meat Pizza" is classified as "Vegetarian" as long as the logic is consistent. The authors advocate for tools like Tawny-OWL, which treats ontology development as programmatic coding, enabling continuous integration and automated unit tests—a massive leap toward professional engineering.
Results and Impact
The impact of rigorous ontology engineering is not just academic. The paper cites external applications where ontologies serve as integrated components:
- Scientific Discovery: The LABORS ontology enabled a "Robot Scientist" to autonomously discover functional genomics knowledge.
- Performance Gains: In gene prioritization tasks, leveraging the hierarchical properties of ontologies led to a 54.1x performance increase over methods relying strictly on raw variant data.
(Note: The modularity of OntoDM allows it to be used across QSAR studies, text mining, and software annotation.)
Critical Insight: The Need for an "IT Portal"
A significant takeaway is the "sociological barrier." While the biomedical community has BioPortal, the IT community lacks a unified repository. This leads to the "Duplication of Effort" problem, where multiple researchers build non-interoperable versions of "Machine Learning Ontologies."
The authors propose that ontologies should be "wrapped" as services within a Service Oriented Architecture (SOA), complete with metadata regarding their inputs, outputs, and quality stages.
Conclusion: Toward a Craft
The transition from "Art" to "Craft" requires the community to:
- Adopt Standardized Testing (reporting test results in journals).
- Build Centralized Repositories for IT-specific knowledge.
- Focus on Objective Comparability to allow different ontologies of the same domain to be measured against one another.
By treating ontologies as software components rather than static documents, we can unlock their potential to drive the next generation of complex, automated information systems.
