Decoding the Law: Rhetorical Structuring in Automatic Legal Summarisation
Extractive summarisation of legal texts
The paper presents a specialized system for the extractive summarisation of legal texts, specifically UK House of Lords judgments. By utilizing a new XML corpus (HOLJ), the authors employ machine learning classifiers to predict the rhetorical status of sentences and rank them for relevance, achieving state-of-the-art accuracy in identifying summary-worthy content.
TL;DR
Processing legal judgments is a Herculean task for both humans and machines due to their length and linguistic complexity. This paper introduces a system that doesn't just "pick sentences," but understands their functional role—whether they are recounting facts (FACT), citing law (BACKGROUND), or delivering a verdict (DISPOSAL). By combining Maximum Entropy sequence modeling with a rich XML-based NLP pipeline, the authors achieve high-precision extraction that can be tailored to different user needs.
Problem & Motivation: Beyond Keyword Matching
In the legal world, a summary that misses the "DISPOSAL" (the final decision) or confuses a "PROCEEDING" (what happened in lower courts) with a "BACKGROUND" (the standing law) is worse than useless—it's misleading.
Standard summarisation techniques (like Lead-based extraction) often fail here because legal judgments are not structured like news articles. The authors realized that to summarize law, you must model the communicative intent of the judge. Their intuition was to treat summarisation as a two-step classification problem:
- What is this sentence doing? (Rhetorical Class)
- Is it important enough to include? (Relevance)
Methodology: The Power of Rhetorical Zoning
The core of the system is the HOLXML pipeline, which transforms raw HTML into a rich linguistic map. Unlike previous models that required massive hand-coded dictionaries of "cue phrases," this system uses a shallow but robust linguistic analysis:
- Named Entity Recognition (NER): Identifying Judges, Acts, and Courts.
- Linguistic Logic: Using verb tense, voice (active/passive), and modality as proxies for rhetorical intent. For example, a "DISPOSAL" sentence often uses present tense active verbs ("I allow the appeal").
Model Architecture
The authors moved beyond simple independent sentence classification by using Sequence Modeling. Since sentences of the same type (like a long list of FACTs) tend to cluster, the model predicts the label of sentence based on the labels of and .

Experiments & Results: Precision is King
The researchers tested various classifiers including C4.5, Naive Bayes, SVM, and Maximum Entropy (ME).
- Rhetorical Performance: The ME sequence model hit an F-score of 61.2. While lower than results in scientific paper domains, the improvement over the baseline (12.0) was actually more significant, proving the model's robustness in a more "noisy" domain.
- Relevance Ranking: The system excelled at finding the "needle in the haystack." In the DISPOSAL category—the most critical part of a legal summary—the model achieved an F-score of 60.3 and incredibly high precision.

The "Disposal" Insight
A critical finding was that ranking alone is insufficient. If you only select the "most relevant" sentences, you might get a summary full of verdicts but missing the crucial "FACT" context. The authors argue for Rhetorical Templates: selecting a specific distribution of sentence types to ensure a balanced, readable summary.
Critical Insight & Conclusion
The real value of this paper isn't just the F-score; it's the Portability. By replacing hand-crafted rules with automatic linguistic features (lemmas, POS tags, and entity types), the authors created a blueprint for legal NLP that can scale across different jurisdictions.
Limitations: The system still struggles with "FACT" sentences because the vocabulary used to describe a crime or a contract is too diverse for simple models to capture without deeper semantic understanding.
The Future: As we move into the era of Large Language Models (LLMs), the "Rhetorical Status" defined here remains a gold standard for how we should evaluate if an LLM-generated summary is legally accurate or just "hallucinating" facts into the disposal.
