Ontology-Driven Data Semantics: Solving the "Ad Hoc" Log Bottleneck in Cyber-Security

Ontology-Driven Data Semantics Discovery for Cyber-Security

2015-01-01
Marcello Balduccini, Sarah Kushner, Jacquelin Speck
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an ontology-driven architecture for data semantics discovery in cyber-security, designed to extract structured information from ad hoc, human-readable log files without predefined formats. By integrating Knowledge Representation with Machine Learning (SVMs), the system automatically identifies file formats, tokenizes data using genetic programming-based regex generation, and maps entities to a comprehensive cyber-asset ontology.

TL;DR

In the high-stakes world of cyber-forensics, the inability to parse unknown or changing log formats (Ad Hoc data) is a critical failure point. This paper presents a hybrid architecture that uses Machine Learning to discover the structure of unknown files and an Ontology to give those structures semantic meaning. The result? A system that can turn a pile of undocumented text files into a searchable, unified knowledge base with ~98% file identification accuracy.

Context: Why "Grep" is No Longer Enough

Security analysts often find themselves in "log hell"—navigating thousands of files from diverse nodes, each with unique configurations. Standard tools like grep or regular-expression-based searches are brittle; they fail the moment a software update changes a log's delimiter or an analyst needs to correlate a "Network Address" that appears as an IPv4 in one file and a MAC address in another. The fundamental challenge is the semantic gap between raw text and actionable intelligence.

Methodology: The Hybrid Intelligence Engine

The paper's architecture is a sophisticated pipeline that bridges machine learning with formal logic.

1. Structure Discovery (The "How")

Instead of manually writing parsers, the system uses a multi-stage ML approach:

  • Novelty Detection: A One-Class SVM determines if a file is "known." If it’s new, it triggers the Template Generator.
  • Genetic Programming: The system identifies delimiters (like whitespace or punctuation) and uses genetic algorithms to "evolve" the perfect regular expression that fits the observed data tokens.
  • Feature Engineering: It uses n-grams of space-delimited tokens, but with a twist—it replaces alphanumeric characters with generic labels (e.g., "192.168.1.1" becomes "NNN.NNN.N.N") to focus on structural patterns rather than specific values.

2. The Ontology (The "Why")

The ontology acts as the brain of the operation. It defines the hierarchy of Events (OS events, hardware events) and Objects (IP addresses, processes).

  • Abstraction: If an analyst searches for a NetworkAddress, the ontology knows to look for IPv4, IPv6, and MACAddress simultaneously.
  • Guided Learning: The ontology stores paths to training samples, allowing the ML components to "self-correct" based on domain knowledge.

Architecture Overview Figure 1: The proposed architecture showing the interaction between ML modules and the domain ontology.

Experiments: Quantifying the Semantic Lift

The researchers tested the system on 2,022 text files. The performance metrics highlight a high degree of precision in high-level categorization:

TaskPrecisionRecallF-Measure
File Format ID0.97910.98080.9799
Record Type ID0.83500.84380.8394
Entity Type ID0.82790.78190.8042

While File Format identification is near-perfect, the slightly lower scores for Entity Type identification (0.80) reflect the inherent ambiguity of short text strings—is "1024" a Port Number, a User ID, or a File Size? However, the authors argue that the Ontology's hierarchy mitigates this: even a misclassified low-level entity might still be correctly grouped under a parent class, preserving query relevance.

Ontology Mapping Example Figure 2: Example of a DNS Query Record being mapped into the hierarchical ontology.

Critical Insight: Semantic Resiliency

The true value of this work lies in Query Independence. Because the data is mapped to a high-level ontology, an analyst can write a SPARQL query to find a "malicious attachment arrival followed by a DNS tunnel" without knowing the specific format of the mail logs or the DNS logs on a particular node. The system handles the "translation" from raw text to conceptual event.

Limitations & Future Work

Despite its strengths, the system relies on the assumption that non-alphanumeric characters serve as delimiters. In modern "structured-within-unstructured" logs (like nested JSON within a syslog), this heuristic might struggle. The authors' plan to explore semi-supervised learning is a necessary next step to reduce the burden of providing labeled training samples for every new file type.

Conclusion

This paper is a significant step toward "Zero-Touch" forensics. By moving away from brittle, human-authored parsers and toward an ontology-driven discovery model, it provides a blueprint for managing the data explosion in large-scale network security.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) to replace or augment the Genetic Programming-based regex generation for ad hoc log parsing.
  • Identify the primary authors and core theories behind the PADS data description language, and analyze how recent "Data Semantics Discovery" frameworks have improved upon its manual specification requirements.
  • Explore how the cyber-asset ontology proposed in this paper compares to newer standardization frameworks like STIX 2.1 or the MITRE ATT&CK knowledge base for event correlation.
Contents
Ontology-Driven Data Semantics: Solving the "Ad Hoc" Log Bottleneck in Cyber-Security
1. TL;DR
2. Context: Why "Grep" is No Longer Enough
3. Methodology: The Hybrid Intelligence Engine
3.1. 1. Structure Discovery (The "How")
3.2. 2. The Ontology (The "Why")
4. Experiments: Quantifying the Semantic Lift
5. Critical Insight: Semantic Resiliency
5.1. Limitations & Future Work
6. Conclusion