NDDO: Bridging the Semantic Gap Between Brain Research and Data Mining

Neurodegenerative Disease Data Ontology

2019-01-01
Ana Kostovska, Ilin Tolovski, Fatima S. Maikore, Larisa N. Soldatova, Pance Panov
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the Neurodegenerative Disease Data Ontology (NDDO), a semantic framework designed to represent clinical, imaging, and biomarker data for brain diseases. It enables interoperability between medical datasets and data mining tools, specifically targeting Alzheimer’s and Parkinson’s disease research within the Human Brain Project (HBP).

TL;DR

The Neurodegenerative Disease Data Ontology (NDDO) is a comprehensive semantic model designed to unify clinical, biomarker, and imaging data for Alzheimer’s and Parkinson’s research. By aligning medical procedures with data mining ontologies (OntoDM), it enables automated dataset annotation and the conversion of ambiguous human-language medical protocols into machine-processable workflows.

Context & Positioning

In the landscape of the Human Brain Project (HBP), the ability to reason over distributed knowledge sources is a bottleneck. NDDO isn't just a vocabulary; it is a structural bridge. It sits at the intersection of Domain Knowledge (Neurodegeneration) and Technical Infrastructure (Data Mining), providing the metadata necessary for "Intelligent Data Mining" systems to understand what they are analyzing.

The Problem: The Clinical-Data Disconnect

Medical data is notoriously messy. A "Visit" in one hospital might imply different assessments than in another, and the protocol for an ELISA assay might be written in ambiguous natural language. For a machine learning (ML) model, this lack of context leads to:

  • Incompatibility: Difficulty in merging datasets from different initiatives (like ADNI and PPMI).
  • Expertise Barrier: Clinicians often cannot build their own ML pipelines because the "data types" don't match the "algorithm requirements" semantically.

Methodology: The Four-Pillar Model

NDDO solves this by modeling the medical domain through a robust hierarchical structure developed in OWL2. It focuses on four central entities that anchor the patient journey:

  1. Study Participant: Distinguishing between the patient and a "study partner" (essential for Alzheimer's where patients may lack self-awareness).
  2. Visit: The temporal link between participants and assessments.
  3. Health Care Process Assay: Categorized into Clinical, Biomarker, and Imaging assessments.
  4. Diagnosis: The logical output derived from assay scores.

NDDO Core Structure

Figure 1: The core class hierarchy of NDDO showing the relationship between participants, visits, and assays.

The "Secret Sauce": Integration with OntoDT and OntoDM

What makes NDDO academically significant is its interoperability. By importing classes from OntoDT (Ontology of Datatypes), it maps medical results (like a cognitive score) directly to computational data types (like an ordered primitive). This allows a reasoner to automatically determine if a dataset is suitable for a specific ML task, such as Multi-Target Regression.

Experimental Use Cases

1. Semantic Annotation for ML

The authors demonstrated that NDDO can semantically label features in the ADNI (Alzheimer's Disease Neuroimaging Initiative) and PPMI (Parkinson's Progression Markers Initiative) datasets.

  • Clustering Task: Annotating unlabeled datasets for grouping male and female patients.
  • Predictive Modeling: Using multi-target regression to predict motor impairment scores from fMRI scans.

Semantic Annotation Schema

Figure 2: Annotation schema for clustering tasks, linking dataset specifications to specific NDDO concepts.

2. Standardizing Laboratory Procedures

By integrating with OCL-SOP (Ontology for Clinical Laboratory Standard Operating Procedures), NDDO converts natural language lab manuals into unambiguous, structured data. The paper shows an example of a Hemoglobin ELISA Kit protocol being translated into a machine-readable table that specifies exact temperatures, volumes, and equipment.

Structured Protocol

Table 1: Example of a machine-processable protocol derived from NDDO and OCL-SOP.

Critical Insight & Future Outlook

The true value of NDDO lies in its Inductive Bias for healthcare AI. By forcing medical data into a logically consistent framework, we move closer to "Self-Describing Data."

Limitations: Currently, the semi-automatic mapping of features to classes requires manual intervention via the Apache Jena library. The next frontier will be the complete automation of this mapping using Large Language Models (LLMs) guided by the NDDO schema.

Takeaway: NDDO is a foundational piece of the puzzle for the Human Brain Project, ensuring that as we collect more brain data, the machine—not just the human—understands what it means.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize the OBO Foundry principles for integrating multi-modal neuroimaging and biomarker data in neurodegenerative disease research.
  • Which paper originally proposed the OntoDM-core (Ontology of Core Data Mining Entities), and how does NDDO extend its class hierarchy for medical diagnosis?
  • Explore current research on using semantic web technologies and ontologies to automate the generation of machine learning pipelines for clinical decision support systems.
Contents
NDDO: Bridging the Semantic Gap Between Brain Research and Data Mining
1. TL;DR
2. Context & Positioning
3. The Problem: The Clinical-Data Disconnect
4. Methodology: The Four-Pillar Model
4.1. The "Secret Sauce": Integration with OntoDT and OntoDM
5. Experimental Use Cases
5.1. 1. Semantic Annotation for ML
5.2. 2. Standardizing Laboratory Procedures
6. Critical Insight & Future Outlook