A Priori Modularisation: Architecting Knowledge for the "Messy" Real World
A priori ontology modularisation in ill-defined domains
The paper introduces an a priori ontology modularisation methodology designed for ill-defined domains. It proposes a three-layered architecture (Upper, High-level Reusable, and Case-specific) to manage cognitively complex knowledge and has been validated within the EU projects Dicode and ImREAL.
TL;DR
In the realm of Semantic Web and Knowledge Engineering, we often try to organize information after it has already become unmanageable. This paper argues for a paradigm shift: A Priori Modularisation. By designing ontologies in three distinct layers (Upper, Reusable, and Case-specific) from the start, researchers can navigate "ill-defined domains"—complex human activities like mentoring or collaborative decision-making—where knowledge is tacit, evolving, and notoriously hard to pin down.
The "Ill-Defined" Challenge: Why Standard Ontologies Fail
Most successful ontologies, like GALEN (Medical) or GENE, operate in "well-defined" domains with rigid hierarchies. However, modern AI and Semantic Web applications are moving toward human-centric processes:
- Tacit Knowledge: Skills like "negotiating" or "motivating" aren't easily codified.
- Expanding Scope: As users interact with social media, new concepts emerge dynamically.
- Heterogeneity: Knowledge comes from psychology, social science, and CS simultaneously.
The authors identify that waiting until an ontology is "finished" to modularize it (a posteriori) leads to massive restructuring costs and poor scalability when dealing with millions of instances.
Methodology: The Three-Layer Blueprint
The core contribution is a structured, top-down layering strategy that anchors abstract theory to specific data.

- Upper Ontology Layer: This is the "theoretical anchor." In the ImREAL project, they used Activity Theory (Subject, Object, Tools, etc.) to define the skeleton of human interaction.
- High-level Reusable Domain Layer: This bridges theory and practice. For example, the abstract concept of "Tools" is specialized into "Mental Tools" (verbal communication) or "Physical Tools" (a CV) that apply across many different scenarios.
- Case-Specific Layer: The "leaves" of the tree. This contains the granular data needed for a specific application, such as a doctor-patient interview or a specific corporate job interview simulation.
The Logic of "A Priori"
Instead of extracting modules from a giant blob of data, the authors use Semantic-driven and Reusability-driven strategies. They decide where a concept belongs based on its generality before it is even coded into the system using tools like ROO (Rabbit to Ontology Authoring).
Real-World Validation: ImREAL and Dicode
The paper demonstrates the methodology across two major EU initiatives:
- ImREAL: Focused on interpersonal communication. By using this layered approach, they could easily adapt a system meant for "Job Interviews" to work for "Cultural Mentoring" simply by swapping the case-specific module while keeping the "Activity Theory" upper layer intact.
- Dicode: Centered on data-intensive decision-making. Here, the ontology helps machines and humans collaborate on medical diagnoses (e.g., Rheumatoid Arthritis) by providing a shared vocabulary for "Sensemaking Operations."
Critical Analysis & Takeaways
The brilliance of this work lies in its recognition that modularity is a design philosophy, not just a technical optimization. While a posteriori tools (like those in the NeOn toolkit) are great for cleaning up old data, they don't help a Knowledge Engineer think through a complex problem from scratch.
Limitations: The paper is a "Short Paper," meaning it lacks deep quantitative benchmarks on reasoning speed improvements (though it suggests modularisation helps). Furthermore, the reliance on high-level frameworks like Activity Theory requires domain experts to be highly skilled in abstract modeling.
Future Outlook
As we move toward Neuro-symbolic AI, where LLMs are grounded by Knowledge Graphs, the "A Priori Modularisation" approach offers a way to keep those graphs manageable. By separating the "logic of activity" from the "specifics of the dataset," we can build AI that is both theoretically sound and practically flexible.
Reference: Thakker, D., et al. (2011). A priori ontology modularisation in ill-defined domains. I-Semantics '11.
