Beyond Rules: Machine Learning for German Ontology-Based Question Classification
Classifying German Questions According to Ontology-Based Answer Types
This paper evaluates three machine learning algorithms (Decision Tree, Naïve Bayes, and k-Nearest Neighbor) to classify German natural language questions into approximately 50 ontology-based answer types. Using a corpus of 1,400 German questions, the research demonstrates that Naïve Bayes combined with shallow features like bag-of-words and statistical collocations provides the most robust classification performance.
TL;DR
Researchers from Saarland University have challenged the traditional rule-heavy approach to German Question Answering (QA). By testing three classic ML algorithms on a corpus of 1,400 questions, they discovered that Naïve Bayes, paired with simple bag-of-words and statistical collocations, outperforms complex manual rules and even Named Entity Recognition (NER). This work provides a pragmatic blueprint for building flexible, ontology-driven QA systems in morphologically complex languages.
The "Rule" Problem in German QA
Historically, analyzing German questions has been a linguistic nightmare. Developers typically relied on hand-coded rules to map a question like "Wann veröffentlichte Milan Füst seine ersten Gedichte?" to an answer type (e.g., Date).
However, rules are brittle. If the domain shifts or the ontology expands, the entire rule-set must be rewritten. The authors identify two core challenges:
- Structural Complexity: German's rich morphology and varied sentence structure make rule-writing exhaustive.
- Ontology Scalability: Moving beyond basic "Who/What/Where" factoids to 50+ hierarchical classes (including abstract "Explanations") makes manual heuristics nearly impossible to maintain.
Methodology: The Power of Shallow Features
The study systematically compares Decision Trees, k-Nearest Neighbor (kNN), and Naïve Bayes. The real North Star of the research, however, isn't just the algorithms—it's the feature engineering.
Key Feature Sets
- Trigger Words: Simple question words (Wann, Wo, Wer).
- Statistical Collocations: Using bigrams that show high standard deviation in occurrence, rather than what "feels" right to a linguist.
- Bag-of-Words (BoW): Treating the question as an unordered set of tokens.
Figure 1: While seemingly simple, the Naïve Bayes model utilizes a conditional probability model (Bayes' Theorem) which effectively handles the sparse nature of the 1,400-question dataset.
Experiments and Counter-Intuitive Results
The authors used 10-fold cross-validation across their corpus. The findings debunk several common assumptions in NLP:
- The "NER" Paradox: Adding Named Entity Recognition labels (Person, Location, Org) actually decreased accuracy in many tests. This is likely due to lack of robustness in NER for short, sentence-level questions.
- The Naïve Bayes Dominance: Despite its "naive" assumption of feature independence, it consistently hit 65% accuracy, outperforming kNN and Decision Trees significantly when combined with statistical collocations.
Table 3: Accuracy comparison showing that the combination of Baseline + Statistical Collocations + Bag-of-Words yields the peak performance.
Critical Insight: Why Shallow Wins
Why did "simple" features beat "complex" linguistic ones?
- Generalization: Hand-picked rules and intuitive collocations are biased by the developer's perspective.
- Redundancy: BoW and statistical collocations often implicitly capture the same information as NER or syntax patterns but with higher coverage and less noise.
Conclusion & Future Directions
The Saarland University team proved that machine learning can effectively bridge the gap between natural language questions and complex ontologies, even in a "difficult" language like German.
Limitations: The system still struggles with "abstract" questions (e.g., "What is the difference between X and Y?") and classes with very few training examples.
The Takeaway for Developers: When building a QA system today, don't start with complex linguistic parsers. Start with a solid statistical baseline. As this 2026-referenced study suggests, the statistical "signals" in the text are often more powerful than the rules we try to impose on them.
Keywords: Question Answering, German NLP, Ontology, Machine Learning, Naïve Bayes, SmartWeb.
