Beyond Keywords: Enhancing Semantic Search with Computer Science Ontology and N-Grams
Semantic Search Using Computer Science Ontology Based on Edge Counting and N-Grams
This paper introduces a semantic search framework for Computer Science literature by combining a domain-specific ontology (based on ACM/IEEE CS2013) with two distinct metrics: Wu and Palmer's edge counting for hierarchical relationships and N-grams for lexical string similarity. The core contribution is a hybrid document similarity score (DSS) that outperforms traditional keyword-based retrieval by addressing synonyms and taxonomic "is-a" relations.
TL;DR
Traditional search engines often fail because they look for words rather than meaning. This paper presents a specialized search framework that utilizes a Computer Science Ontology alongside N-gram string matching to bridge the semantic gap. By calculating the "distance" between concepts in a taxonomic hierarchy (based on ACM/IEEE standards), the authors provide a more nuanced ranking system for academic documents.
The "Vocabulary Mismatch" Problem
In a standard keyword search, if you search for "Machine Learning," a document titled "Neural Networks" might be overlooked if the specific keyword isn't present. This is the synonym/hyponym problem.
Existing tools like WordNet are too general for technical fields. For instance, the hierarchical relationship between "Decision Making" and "Software Project Management" is critical in a CS context but often absent in general-purpose lexicons. Consequently, researchers face "information overload"—thousands of results, many irrelevant, while the most pertinent ones remain buried.
Methodology: The Hybrid Scoring Engine
The authors solve this by introducing a multi-tiered weighting system. Instead of binary matching (Found/Not Found), they use a Semantic Weight Keyword (SWK) approach.
1. The Computer Science Ontology (Taxonomic Backbone)
The researchers built an ontology based on the CS2013 curricula (ACM/IEEE). This ontology organizes 18 Knowledge Areas (KAs) such as Intelligent Systems and Software Engineering across four levels.
To calculate the weight between terms, they use the Wu and Palmer Measure:
- Physical Intuition: The closer two nodes are in the "family tree" of Computer Science, the higher their similarity score.
- If keywords share the same sub-area, they receive a weight of 0.75; if they are in the same general area, 0.5.
2. N-Grams (The Lexical Safety Net)
Technical papers often use idiosyncratic terminology. The Tri-gram method breaks keywords into 3-character sequences to find string-level similarities even when terms aren't in the ontology.
3. Calculating the Document Similarity Score (DSS)
The final score for a document is an aggregation of these weights:
SWK = SWK(n-gram) + SWK(cs-onto)

Fig 1: Illustration of the scoring process where Query Keywords are evaluated against Document Keywords via the CS Ontology.
Experimental Validation
Using a dataset of 1,769 documents, the authors tested the system with "Decision Making" as a query.
Key Findings:
- Effective Ranking: Documents like Paper #115 (scoring 1.493) were correctly ranked highest because they contained both the exact query term and related ontological concepts (Architecture, Design Quality).
- Statistical Significance: A t-test comparison with other established weights (Hliaoutakis and Nicola) showed a p-value of 0.105 (> 0.05), confirming that the proposed weighting method is consistent with current SOTA semantic approaches while being specifically optimized for CS.

Table 1: Document Similarity Scores for the query "Decision Making" across different Knowledge Areas.
Critical Insight & Conclusion
While the paper demonstrates an effective way to handle domain-specific search, its reliance on a static ontology (CS2013) is a potential limitation. As Computer Science evolves (e.g., the explosion of Generative AI), ontologies must be dynamically updated.
Takeaway: This work demonstrates that for professional domains, ontological structure is the "anchor" while lexical matching (N-grams) is the "flexible cord"—combining them creates a retrieval system that understands context better than any keyword list ever could.
Future Outlook
The authors suggest that future iterations will involve Information Visualization, allowing users to see the "distance" between their query and the results on a semantic map, further reducing the cognitive load on researchers.
