Beyond Bag-of-Words: Leveraging Tree Kernels and Discourse Theory for Portuguese Text Classification

Text Classification Using Tree Kernels and Linguistic Information

2008-01-01
Teresa Gonçalves, Paulo Quaresma
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores integrating syntactic and semantic structures into text classification for the Portuguese language using SVMs and Tree Kernels. It introduces a novel discourse-structure representation based on Discourse Representation Theory (DRT) that achieves classification performance comparable to the traditional Bag-of-Words (BoW) model.

TL;DR

For decades, the "Bag-of-Words" (BoW) model has been the workhorse of text classification, but it remains "blind" to the underlying logic of language. This research challenges the status quo by introducing Tree Kernels applied to Discourse Representation Structures (DRS). While syntax alone proves noisy, a structured semantic approach matches traditional performance while slashing feature complexity by 30%.

Background: The Limits of Word Statistics

In the context of morphologically rich languages like Portuguese, a single verb can have 66 different forms. Standard Machine Learning often struggles unless significant pre-processing (like lemmatization) is applied. However, even with lemmatization, the relationship between entities—who did what to whom—is lost. The authors investigate whether "Deep" linguistic structures (syntax and semantics) can provide a more robust signal for SVM classifiers than simple word counts.

Problem & Motivation: Why Structure Usually Fails

Historically, adding syntactic information to text classifiers hasn't helped much. The "Noise" problem is the primary culprit: parse trees are highly specific, and the variation in sentence structure across a 4,000-document dataset often leads to data sparsity. The authors hypothesize that while syntax might be too granular, semantics (the logical form of a document) might capture the "essence" of a topic more cleanly.

Methodology: Mapping Meaning to Trees

The core innovation lies in how the authors represent a document's meaning. Instead of a flat vector, they use:

  1. Syntactic Trees: Based on the PALAVRAS parser, representing sentences as ordered constituent trees.
  2. Discourse-Structure Representation: This is the "secret sauce." Using Discourse Representation Theory (DRT), they create a logical form of every sentence.
  3. Referent Substitution: To make these structures useful for a kernel, they replace abstract discourse referents (like ) with concrete names and properties.

Concept of Kernel Mapping Figure 1: The Kernel function maps non-linear linguistic patterns into a high-dimensional feature space where linear separation is possible.

The researchers used a Convolution Tree Kernel, which calculates document similarity by counting common subtrees. This allows the SVM to "see" recurring logical patterns across different texts.

Experiments & Results: Semantics vs. Statistics

The study utilized the Público95 dataset (Portuguese newspaper articles). The results yielded a fascinating dichotomy:

  • The Syntactic Failure: Directly using parse trees (tree.tot) lagged behind the Bag-of-Words baseline (F1 0.812 vs 0.857). This confirms that raw syntax is often too specific for general topic classification.
  • The Semantic Success: The discourse-structure (dis.noun+pro) achieved an F1-micro of 0.833. While numerically slightly lower than BoW, statistical significance tests showed they are effectively equivalent in discriminative power.

Semantic Tree Transformation Figure 2: Workflow from original document to SIN2SEM output and the final discourse-structure representation.

The "Hidden" Benefit: Efficiency

Perhaps the most impressive result isn't the accuracy, but the feature reduction. The semantic model used only 46,186 types compared to the 70,743 types required by the BoW approach. By focusing on logical conditions rather than every literal word form, the model achieved a 30% reduction in dimensionality without sacrificing performance.

Critical Analysis & Conclusion

Takeaway

This work proves that semantic information is a viable "attribute selector." By distilling a document into its logical entities and their relationships, we can represent its "topic" more efficiently than word frequencies alone.

Limitations

  1. Parser Dependency: The system's quality is bottlenecked by the PALAVRAS and SIN2SEM tools. Errors in the initial parse propagate through to the semantic tree.
  2. Domain Specificity: News articles have a very distinct structure. It remains to be seen if discourse theory holds up as well in informal text (e.g., social media).

Future Outlook

The authors suggest that as Natural Language Processing tools for anaphora resolution and named entity recognition improve, the semantic representations will become even more "pure," likely surpassing the performance of statistical models. In an era where we often prioritize "bigger" models, this paper reminds us that "smarter" representations can still lead to significant gains.

Find Similar Papers

Try Our Examples

  • Which recent papers have successfully combined Tree Kernels with Deep Learning architectures for text classification tasks?
  • What is the origin of the "Subset Tree Kernel" proposed by Collins and Duffy, and how has it been optimized for large-scale datasets since 2002?
  • Are there any studies applying Discourse Representation Theory (DRT) to modern Transformer-based models for long-document understanding?
Contents
Beyond Bag-of-Words: Leveraging Tree Kernels and Discourse Theory for Portuguese Text Classification
1. TL;DR
2. Background: The Limits of Word Statistics
3. Problem & Motivation: Why Structure Usually Fails
4. Methodology: Mapping Meaning to Trees
5. Experiments & Results: Semantics vs. Statistics
5.1. The "Hidden" Benefit: Efficiency
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook