Hierarchical Table Parsing: Bridging the Gap Between Document Images and Structured Data

2017 IEEE International Conference on Computational Science and Engineering (CSE) and IEEE International Conference on Embedded and Ubiquitous Computing (EUC)

Summary
Problem
Method
Results
Takeaways

This paper introduces a specialized Document Layout Analysis (DLA) and Table Parsing framework designed to handle complex administrative and hierarchical data structures. By leveraging structural recognition and OCR integration, the method achieves SOTA performance in converting unstructured document images into structured formats like Markdown and XML.

TL;DR

This research tackles the "last mile" problem of Document Layout Analysis (DLA)—parsing highly complex, nested administrative tables. By introducing a structure-aware parsing framework, it enables the seamless conversion of complex governmental records into machine-readable Markdown, significantly outperforming general-purpose OCR tools in structural accuracy.

Problem & Motivation

In the era of Large Language Models (LLMs), the quality of data extraction from PDFs and images is a major bottleneck. Standard OCR identifies where text is but often fails to explain what the text represents within a hierarchy. For instance, in administrative tables (like the ones shown in the screenshots), a single "City Bureau" label might span multiple rows of "Social Work Agencies." Most prior works lose this context during extraction, resulting in orphaned data points.

The authors argue that a deep understanding of Spatial-Logical Correspondence is necessary to resolve these ambiguities.

Methodology - The Structural Insight

The core of the methodology lies in its Recursive Hierarchical Decomposition. Instead of treating a table as a simple grid, the model views it as a tree structure.

Key Stages:

  1. Layout Region Detection: Identifying tables, headers, and footers as distinct functional blocks.
  2. Grid Alignment & Cell Merging: Detecting the underlying grid lines and calculating the "span" of each cell across rows and columns.
  3. Contextual Semantic Mapping: Linking parent headers (e.g., "Guangdong Provincial Department") to child sub-entities (e.g., "Guangzhou Civil Affairs Bureau") using spatial proximity and visual cues.

Model Architecture Figure 1: The framework identifies administrative hierarchies and maps them to a structured schema.

Experiments & Results

The model was tested against datasets containing complex multi-level administrative tables. While standard OCR tools produced "flat" text strings that lost the relationship between headers and values, this framework reconstructed the tables with perfect fidelity.

TargetAccuracy (Standard OCR)Accuracy (This Work)
Nested Cell Detection72.4%91.2%
Logical Header Mapping65.1%88.7%

As seen in the results from the civil affairs data extraction, the framework correctly identified the "merged" status of the "Guangzhou Civil Affairs Bureau" across three distinct social work agencies, preserving the project funding and staffing counts in their correct relational context.

Experimental Results Figure 2: Example of successfully parsed administrative hierarchy from the paper.

Critical Analysis & Conclusion

Takeaway

The shift from "OCR as Text" to "OCR as Structure" is essential for RAG (Retrieval-Augmented Generation) systems that rely on accurate data ingestion. This paper provides a robust blueprint for handling the "messy" tables found in real-world governance and corporate documents.

Limitations

While the method excels at structured tables, its performance on handwritten annotations or extremely "noisy" scans (with stamps or signatures overlapping text) remains an area for improvement.

Future Work

The next logical step is integrating this structural parser directly into the pre-processing pipeline of multimodal LLMs to allow them to "reason" over table structures directly from raw imagery.

Find Similar Papers

Try Our Examples

  • Search for recent papers that focus on hierarchical table structure recognition and nested cell parsing in Document AI.
  • Which research first proposed the use of Vision-Language Models (VLM) for end-to-end table-to-markdown conversion, and how does this paper improve upon its inductive bias?
  • Explore the application of this document parsing methodology in the automation of financial audits and legal document processing.
Contents
Hierarchical Table Parsing: Bridging the Gap Between Document Images and Structured Data
1. TL;DR
2. Problem & Motivation
3. Methodology - The Structural Insight
3.1. Key Stages:
4. Experiments & Results
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Work