TextTiling: Deciphering the Hidden Structure of Multi-Paragraph Discourse
Multi-Paragraph Segmentation of Expository Text
The paper introduces TextTiling, a domain-independent algorithm designed to segment long, expository texts into multi-paragraph discourse units. By analyzing lexical frequency and distribution, specifically term repetition, it successfully identifies subtopic transitions with performance matching human judgment trends.
TL;DR
How do you navigate a massive wall of text without subheadings? TextTiling is a classic but foundational algorithm that partitions long expository documents into "tiles" or subtopics. By measuring the "lexical flow"—how vocabulary shifts from one paragraph to the next—it identifies where one discussion ends and another begins, achieving accuracy levels comparable to human readers.
Background: The Motivation for Segmentation
In the world of Information Retrieval (IR), size matters. A whole document is often too broad to be a useful search result, while a single paragraph is too narrow. The "sweet spot" is the subtopic—a coherent multi-paragraph unit. Marti Hearst's work addresses the problem of expository texts (like science articles) that lack structural demarcation. Unlike narrative texts that rely on characters or time shifts, expository texts are driven by shifts in subject matter themes.
The Problem: Why Simple Word Chains Fail
Earlier models attempted to use "lexical chains"—tracking a single word's reappearance throughout a text. However, as Hearst points out, long texts are messy. Multiple themes overlap simultaneously. A discussion about the "Moon" might overlap with "planets," "tides," and "gravity" all at once. Relying on a single chain to break isn't enough; you need to find where the bulk of these threads change in a "maximal way."
Methodology: The Geometry of Similarity
The TextTiling algorithm (specifically the Block Similarity version) operates on three main stages:
1. Tokenization and Blocking
The text is divided into "token-sequences" (usually 20 words). These are then grouped into "blocks" (size , usually the average paragraph length).
2. Similarity Computation
Using a sliding window, the algorithm compares two adjacent blocks ( and ). It calculates the Cosine Similarity based on term frequency:
High similarity implies the subtopic is continuing; a sudden drop (a valley) suggests a potential boundary.
3. Boundary Identification via Depth Scores
Instead of using absolute thresholds, Hearst introduces Depth Scores. It looks at a valley and measures how much the similarity "rises" on both the left and right sides. This captures the relative change in discourse, making it robust against different writing styles.
Figure 1: Conceptual graph showing sentence connectivity. TextTiling aims for the 'Piecewise Monolithic Structure'—blocks of high connectivity linked sequentially.
Experimental Results: Machines vs. Humans
The algorithm was tested against seven human judges across 13 magazine articles. The results were striking:
- Performance: TextTiling achieved a Precision of 0.66, far exceeding the random baseline of 0.43.
- Refinement: If the algorithm is allowed to be off by just one paragraph (a common case where a summary paragraph confuses the signals), its precision jumps to 0.83.
Figure 2: The similarity plot for the text "Stargazers." The vertical lines represent boundaries identified by the algorithm, aligning closely with where human readers perceived topic shifts.
Critical Analysis & Future Outlook
TextTiling's Strength is its simplicity. It is domain-independent and doesn't require a complex knowledge base or a thesaurus (in fact, Hearst found that adding thesaural information sometimes degraded performance due to noise).
Limitations:
- Sensitivity to Summaries: Paragraphs that summarize previous points (using the same vocabulary) can "trick" the algorithm into seeing high similarity where a boundary should actually exist.
- Linear Assumption: The model assumes subtopics are a flat sequence, ignoring hierarchical structures where one subtopic might be nested within another.
Why it matters today: In the era of LLMs and Retrieval-Augmented Generation (RAG), segmenting long documents into semantically coherent "chunks" is more critical than ever. TextTiling provides a mathematically grounded alternative to "fixed-sized chunking," ensuring that context is preserved within its natural subtopical boundaries.
Conclusion
TextTiling remains a landmark in NLP because it proved that lexical distribution alone could mirror human cognitive segments in discourse. It shifted the focus from "what is the word" to "how do the words flow together."
