Bridging the Legal Gap: Lexical-Morphological Modeling in Legal QA
Lexical-Morphological Modeling for Legal Text Analysis
This paper introduces a hybrid lexical-morphological framework for Legal Information Retrieval (IR) and Question Answering (QA), specifically targeting the COLIEE competition. The core method, R2NC (Ranking Related N-gram Collections), utilizes a mixed-size n-gram model and an AdaBoost ensemble classifier with Word2Vec embeddings to achieve 2nd place in the IR task and 3rd place in the combined IR+QA task.
TL;DR
Answering legal questions is notoriously difficult due to abstract terminology and the high cost of expert knowledge. This paper presents a robust, data-efficient method using R2NC (a mixed-size n-gram ranking model) and AdaBoost classification with Word2Vec embeddings. It effectively handles the "low lexical overlap" problem in legal texts, securing top-3 rankings in the COLIEE competition.
Background: The "Information Crisis" in Law
The legal domain is currently facing an information crisis. While legal documents are more accessible than ever, our ability to analyze them hasn't kept pace. Existing systems often rely on manual Knowledge Engineering (ontologies), which are expensive and brittle. The core challenge lies in Abstraction: a question about "preventing a crime" might map to an article about "management of business without obligation." Traditional keyword searches simply won't find that link.
Methodology: R2NC and Semantic Entailment
The authors propose a two-phase architecture designed to move from raw text to a final YES/NO answer.
1. Relevance Analysis (R2NC)
Instead of standard TF-IDF, the authors developed R2NC (Ranking Related N-gram Collections).
- Morphological Focus: It uses lemmatization and mixed-size n-grams (1 to 3 words) to maintain consistency across topics.
- Length Normalization: Recognizing that law articles are significantly longer than questions, they introduced specific weights ( and ) into the scoring formula to prevent long articles from drowning out relevant short ones.

2. Textual Entailment (TE) via AdaBoost
Once relevant articles are found, the system must decide if they actually "entail" the answer.
- Feature Engineering: The model uses 15 features across distance-based metrics (Euclidean, Cosine) and statistical metrics (Word Overlap, TF-IDF).
- The Word2Vec Edge: By training Word2Vec on the Japanese Civil Code, the model captures semantic neighbors. For instance, it learns that "fees" is semantically close to "costs," allowing it to solve the "H18-2-4" case where lexical overlap is minimal.
Performance & Experiments
The results substantiate that simplicity, when paired with the right features, can outperform complexity.
- Information Retrieval: R2NC achieved an F-measure of 0.54. In "Top 3" recall, it hit 0.64, significantly higher than the TPP baseline of 0.52.
- Classification: AdaBoost-DecisionStump achieved 61.42% accuracy, outperforming standard SVMs by nearly 7%.

Critical Insights: Why It Works (and Where it Fails)
The paper’s success stems from its Inductive Bias—the choice to emphasize morphological lemmas over raw surface forms. This is particularly effective in legal Japanese/English where verb conjugations and specific noun-forms carry heavy semantic weight.
The Overfitting Trap: Despite high training accuracy (61%), the system's performance dropped during the actual competition phase two (37.88%). The authors candidly point to overfitting. Since legal datasets are small (267 pairs), the classifier likely "memorized" certain phrasing patterns that didn't generalize to the test set.
Conclusion & Future Outlook
This work demonstrates that for niche, high-abstraction domains:
- Contextual Similarity (Word2Vec/Sent2Vec) is non-negotiable for bridging the vocabulary gap.
- Structural Analysis (like decomposing law into "requisites" and "effectuations") is the next frontier.
Future systems should look toward combining these lexical-morphological foundations with modern LLMs to better handle the second level of entailment: negation and antonym detection, which this study identified as a current limitation.
