Beyond Bag-of-Words: Leveraging Tree Kernels and Discourse Theory for Portuguese Text Classification
Text Classification Using Tree Kernels and Linguistic Information
This paper explores integrating syntactic and semantic structures into text classification for the Portuguese language using SVMs and Tree Kernels. It introduces a novel discourse-structure representation based on Discourse Representation Theory (DRT) that achieves classification performance comparable to the traditional Bag-of-Words (BoW) model.
TL;DR
For decades, the "Bag-of-Words" (BoW) model has been the workhorse of text classification, but it remains "blind" to the underlying logic of language. This research challenges the status quo by introducing Tree Kernels applied to Discourse Representation Structures (DRS). While syntax alone proves noisy, a structured semantic approach matches traditional performance while slashing feature complexity by 30%.
Background: The Limits of Word Statistics
In the context of morphologically rich languages like Portuguese, a single verb can have 66 different forms. Standard Machine Learning often struggles unless significant pre-processing (like lemmatization) is applied. However, even with lemmatization, the relationship between entities—who did what to whom—is lost. The authors investigate whether "Deep" linguistic structures (syntax and semantics) can provide a more robust signal for SVM classifiers than simple word counts.
Problem & Motivation: Why Structure Usually Fails
Historically, adding syntactic information to text classifiers hasn't helped much. The "Noise" problem is the primary culprit: parse trees are highly specific, and the variation in sentence structure across a 4,000-document dataset often leads to data sparsity. The authors hypothesize that while syntax might be too granular, semantics (the logical form of a document) might capture the "essence" of a topic more cleanly.
Methodology: Mapping Meaning to Trees
The core innovation lies in how the authors represent a document's meaning. Instead of a flat vector, they use:
- Syntactic Trees: Based on the PALAVRAS parser, representing sentences as ordered constituent trees.
- Discourse-Structure Representation: This is the "secret sauce." Using Discourse Representation Theory (DRT), they create a logical form of every sentence.
- Referent Substitution: To make these structures useful for a kernel, they replace abstract discourse referents (like ) with concrete names and properties.
Figure 1: The Kernel function maps non-linear linguistic patterns into a high-dimensional feature space where linear separation is possible.
The researchers used a Convolution Tree Kernel, which calculates document similarity by counting common subtrees. This allows the SVM to "see" recurring logical patterns across different texts.
Experiments & Results: Semantics vs. Statistics
The study utilized the Público95 dataset (Portuguese newspaper articles). The results yielded a fascinating dichotomy:
- The Syntactic Failure: Directly using parse trees (tree.tot) lagged behind the Bag-of-Words baseline (F1 0.812 vs 0.857). This confirms that raw syntax is often too specific for general topic classification.
- The Semantic Success: The discourse-structure (dis.noun+pro) achieved an F1-micro of 0.833. While numerically slightly lower than BoW, statistical significance tests showed they are effectively equivalent in discriminative power.
Figure 2: Workflow from original document to SIN2SEM output and the final discourse-structure representation.
The "Hidden" Benefit: Efficiency
Perhaps the most impressive result isn't the accuracy, but the feature reduction. The semantic model used only 46,186 types compared to the 70,743 types required by the BoW approach. By focusing on logical conditions rather than every literal word form, the model achieved a 30% reduction in dimensionality without sacrificing performance.
Critical Analysis & Conclusion
Takeaway
This work proves that semantic information is a viable "attribute selector." By distilling a document into its logical entities and their relationships, we can represent its "topic" more efficiently than word frequencies alone.
Limitations
- Parser Dependency: The system's quality is bottlenecked by the PALAVRAS and SIN2SEM tools. Errors in the initial parse propagate through to the semantic tree.
- Domain Specificity: News articles have a very distinct structure. It remains to be seen if discourse theory holds up as well in informal text (e.g., social media).
Future Outlook
The authors suggest that as Natural Language Processing tools for anaphora resolution and named entity recognition improve, the semantic representations will become even more "pure," likely surpassing the performance of statistical models. In an era where we often prioritize "bigger" models, this paper reminds us that "smarter" representations can still lead to significant gains.
