Beyond Rules: Machine Learning for German Ontology-Based Question Classification

Classifying German Questions According to Ontology-Based Answer Types

2007-01-01
Adriana Davidescu, Andrea Heyl, Stefan Kazalski, Irene M. Cramer, Dietrich Klakow
Summary
Problem
Method
Results
Takeaways
Abstract

This paper evaluates three machine learning algorithms (Decision Tree, Naïve Bayes, and k-Nearest Neighbor) to classify German natural language questions into approximately 50 ontology-based answer types. Using a corpus of 1,400 German questions, the research demonstrates that Naïve Bayes combined with shallow features like bag-of-words and statistical collocations provides the most robust classification performance.

TL;DR

Researchers from Saarland University have challenged the traditional rule-heavy approach to German Question Answering (QA). By testing three classic ML algorithms on a corpus of 1,400 questions, they discovered that Naïve Bayes, paired with simple bag-of-words and statistical collocations, outperforms complex manual rules and even Named Entity Recognition (NER). This work provides a pragmatic blueprint for building flexible, ontology-driven QA systems in morphologically complex languages.

The "Rule" Problem in German QA

Historically, analyzing German questions has been a linguistic nightmare. Developers typically relied on hand-coded rules to map a question like "Wann veröffentlichte Milan Füst seine ersten Gedichte?" to an answer type (e.g., Date).

However, rules are brittle. If the domain shifts or the ontology expands, the entire rule-set must be rewritten. The authors identify two core challenges:

  1. Structural Complexity: German's rich morphology and varied sentence structure make rule-writing exhaustive.
  2. Ontology Scalability: Moving beyond basic "Who/What/Where" factoids to 50+ hierarchical classes (including abstract "Explanations") makes manual heuristics nearly impossible to maintain.

Methodology: The Power of Shallow Features

The study systematically compares Decision Trees, k-Nearest Neighbor (kNN), and Naïve Bayes. The real North Star of the research, however, isn't just the algorithms—it's the feature engineering.

Key Feature Sets

  • Trigger Words: Simple question words (Wann, Wo, Wer).
  • Statistical Collocations: Using bigrams that show high standard deviation in occurrence, rather than what "feels" right to a linguist.
  • Bag-of-Words (BoW): Treating the question as an unordered set of tokens.

Model Comparison Overview Figure 1: While seemingly simple, the Naïve Bayes model utilizes a conditional probability model (Bayes' Theorem) which effectively handles the sparse nature of the 1,400-question dataset.

Experiments and Counter-Intuitive Results

The authors used 10-fold cross-validation across their corpus. The findings debunk several common assumptions in NLP:

  1. The "NER" Paradox: Adding Named Entity Recognition labels (Person, Location, Org) actually decreased accuracy in many tests. This is likely due to lack of robustness in NER for short, sentence-level questions.
  2. The Naïve Bayes Dominance: Despite its "naive" assumption of feature independence, it consistently hit 65% accuracy, outperforming kNN and Decision Trees significantly when combined with statistical collocations.

Experimental Results Table Table 3: Accuracy comparison showing that the combination of Baseline + Statistical Collocations + Bag-of-Words yields the peak performance.

Critical Insight: Why Shallow Wins

Why did "simple" features beat "complex" linguistic ones?

  • Generalization: Hand-picked rules and intuitive collocations are biased by the developer's perspective.
  • Redundancy: BoW and statistical collocations often implicitly capture the same information as NER or syntax patterns but with higher coverage and less noise.

Conclusion & Future Directions

The Saarland University team proved that machine learning can effectively bridge the gap between natural language questions and complex ontologies, even in a "difficult" language like German.

Limitations: The system still struggles with "abstract" questions (e.g., "What is the difference between X and Y?") and classes with very few training examples.

The Takeaway for Developers: When building a QA system today, don't start with complex linguistic parsers. Start with a solid statistical baseline. As this 2026-referenced study suggests, the statistical "signals" in the text are often more powerful than the rules we try to impose on them.


Keywords: Question Answering, German NLP, Ontology, Machine Learning, Naïve Bayes, SmartWeb.

Find Similar Papers

Try Our Examples

  • Find recent papers on German question classification that utilize Deep Learning or Transformer-based models to compare with traditional ML benchmarks.
  • What is the origin of the SmartWeb ontology and how has it evolved for multimodal question answering since Sonntag et al. (2006)?
  • Explore research applying statistical collocation and bag-of-words features to question answering tasks in other morphologically rich languages like Finnish or Turkish.
Contents
Beyond Rules: Machine Learning for German Ontology-Based Question Classification
1. TL;DR
2. The "Rule" Problem in German QA
3. Methodology: The Power of Shallow Features
3.1. Key Feature Sets
4. Experiments and Counter-Intuitive Results
5. Critical Insight: Why Shallow Wins
6. Conclusion & Future Directions