OntoDT: Building a Semantic Foundation for Global Datatype Interoperability

Generic ontology of datatypes

2015-08-13
Pance Panov, Larisa N. Soldatova, Saso Dzeroski
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces OntoDT, a generic ontology designed for the formal representation of datatypes in scientific knowledge. Based on the ISO/IEC 11404 standard, it provides a machine-processable framework for defining datatype properties, taxonomies, and characterizing operations, achieving state-of-the-art interoperability for data mining and bioinformatics.

TL;DR

OntoDT is a comprehensive, open-source ontology designed to give "meaning" to datatypes. By formalizing entities like "integers," "records," and "sequences" based on international standards (ISO/IEC 11404), it enables different scientific software and dataset repositories to communicate seamlessly. It provides a rigorous taxonomy that moves beyond simple labels to define what operations (like addition or comparison) are logically valid for specific data.

Problem & Motivation: The "Tower of Babel" in Data Science

In scientific research, we often treat datatypes as self-evident. However, for a computer, the difference between a "Date," a "Timestamp," and a "Project Start" is often lost without explicit semantic mapping.

The authors identify two major pain points:

  1. Semantic Ambiguity: Existing standards like XSD (XML Schema Definition) are widely used but lack a formal logical model. This leads to users creating redundant custom types for the same underlying concepts.
  2. Workflow Fragmentation: In data mining, connecting a preprocessing tool to a learning algorithm is manual and error-prone because the "input/output" requirements are not machine-readable at a deep level.

The Insight: By building an ontology on top of the Information Artifact Ontology (IAO), we can represent datatypes not just as storage containers, but as "Information Content Entities" with specific mathematical qualities and valid operations.

Methodology: The Architecture of a Datatype

OntoDT isn't just a list; it’s a relational graph. The core methodology splits a datatype into four critical components:

  • Value Space: What values are allowed? (e.g., all integers from to ).
  • Characterizing Operations: What can you do with it? (e.g., a "monadic" operation like Negate or a "dyadic" one like Add).
  • Datatype Properties (Qualities): Is it ordered? Is it numeric? Is it bounded?
  • Datatype Generators: How do we build complex types (like Arrays or Records) from primitive ones?

OntoDT Core Architecture Figure 1: The structural representation of a Datatype class in OntoDT, highlighting the relations between value spaces, operations, and qualities.

One of the most powerful features is the Generator system. For example, a "Positive Integer" isn't a new primitive; it's an "Extended Datatype" created by applying a range generator to the base integer datatype.

The Datatype Taxonomy

The paper proposes a dual-branch taxonomy that organizes data into:

  1. Primitive Datatypes: Atomic entities like boolean, integer, or date-and-time.
  2. Generated Datatypes: Complex structures built via generators, subdivided into Aggregate (e.g., Sets, Records) and Non-aggregate (e.g., Pointers).

Datatype Taxonomy Figure 2: The hierarchical classification of datatypes within OntoDT.

Experiments & Real-World Use Cases

The authors didn't just build a theory; they validated it across three domains:

1. Data Mining (OntoDM-core)

By mapping datasets to OntoDT types, the authors created a Taxonomy of Datasets. This allows a system to automatically determine if an algorithm (e.g., Random Forest) is compatible with a specific dataset (e.g., Multi-label Classification) based on the target datatype's qualities.

2. Dataset Repositories (UCI)

The team annotated 193 datasets from the UCI Machine Learning Repository. They demonstrated that by using a reasoner (HermiT), they could perform complex semantic searches, such as: "Find all datasets where the target output is a homogenous aggregate with key-based access."

3. Bioinformatics (BioXSD)

Biological data is notoriously complex. OntoDT was used to refine BioXSD, linking nucleotide sequences directly to the ChEBI chemical ontology, ensuring that "data about DNA" is logically tied to the "chemical reality of DNA."

Bioinformatics Extension Figure 3: Extending OntoDT to represent domain-specific bio-sequence records.

Critical Insight & Conclusion

OntoDT represents a "mid-level" breakthrough. While upper-level ontologies are too abstract and domain ontologies are too specific, OntoDT provides the shared language of computation.

Takeaway: For the future of AI and automated science, we need more than just "Big Data"; we need "Smart Data" that carries its own structural logic. OntoDT is a significant step toward a world where software can automatically configure itself to process any scientific dataset it encounters.

Limitations: The current version focuses heavily on traditional programming types. Future iterations will need to incorporate high-dimensional types like Tensors, Audio, and Video stream dynamics to remain relevant in the age of Deep Learning.

Find Similar Papers

Try Our Examples

  • Find recent papers published after 2015 that extend the OntoDT ontology or integrate it into newer data mining frameworks like OpenML.
  • What are the original definitions of datatypes in the ISO/IEC 11404:2007 standard, and how does OntoDT map these to the Basic Formal Ontology (BFO)?
  • Explore how OntoDT or similar datatype ontologies are being utilized to improve the orchestration of web services in modern cloud computing environments.
Contents
OntoDT: Building a Semantic Foundation for Global Datatype Interoperability
1. TL;DR
2. Problem & Motivation: The "Tower of Babel" in Data Science
3. Methodology: The Architecture of a Datatype
4. The Datatype Taxonomy
5. Experiments & Real-World Use Cases
5.1. 1. Data Mining (OntoDM-core)
5.2. 2. Dataset Repositories (UCI)
5.3. 3. Bioinformatics (BioXSD)
6. Critical Insight & Conclusion