CNN vs. SVM: Scaling Predictive Coding in the Legal Frontier

Empirical Study of Deep Learning for Text Classification in Legal Document Review

2018-12-01
Fusheng Wei, Han Qin, Shi Ye, Haozhen Zhao
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an empirical study evaluating Deep Learning, specifically Convolutional Neural Networks (CNN), for text classification in legal document review (Predictive Coding). Comparing CNN against the industry-standard Support Vector Machines (SVM) across four real-world legal datasets, the study demonstrates that CNNs achieve superior precision and stability, particularly as training data volume increases.

TL;DR

In the high-stakes world of legal document review—where missing a single "responsive" document can cost millions—the industry has long favored the reliability of Support Vector Machines (SVM). This empirical study by Ankura researchers challenges the status quo, proving that Convolutional Neural Networks (CNN) significantly outperform traditional methods in precision and stability when supplied with sufficient data.

Contextual Motivation

The legal industry faces a data deluge. "Predictive Coding" or Technology Assisted Review (TAR) is no longer a luxury but a necessity. Historically, linear models (SVM/LR) reigned supreme due to their speed and simplicity. However, they are inherently limited by their "Bag-of-Words" nature, ignoring the syntax and sequence that often define legal context. The researchers sought to determine if the feature-extraction prowess of CNNs, which revolutionized image and sentiment analysis, could be effectively "transplanted" into the rigid requirements of legal discovery.

The Architecture of Legal Insight

Unlike traditional models that treat a document as a set of independent word counts, the proposed CNN architecture views legal text as a temporal sequence.

Key Technical Components:

  1. Sequence-Aware Embeddings: Words are mapped to 100-dimensional vectors. Interestingly, the study found that self-trained embeddings outperformed pre-trained GloVe vectors, suggesting that the "legal dialect" is distinct enough that generalized embeddings may dilute accuracy.
  2. 1D Convolutional Layers: These act as "feature detectors" for legal idioms and phrases, scanning the text for local clues that signify relevance.
  3. Global Max Pooling: This reduces the high-dimensional convolutional output to the most salient signals, effectively "flagging" the most relevant parts of a document.

Model Architecture Table 1: The CNN architecture utilized, showing the flow from Embedding to Dense output.

Experimental Battleground

The researchers tested their hypothesis on four distinct real-world projects (A, B, C, D), each containing millions of records. They created "learning curves" by incrementalizing the training sets to observe how model maturity affects performance.

Key Findings:

  • Data Scarcity vs. Abundance: On very small datasets, traditional SVM sometimes held its own. However, as the volume grew, the CNN's accuracy and precision pulled ahead convincingly.
  • The Precision Edge: In the legal world, "Precision at a specific Recall" (e.g., 75%) is the gold standard. In Project D, CNN demonstrated a clear lead, suggesting fewer "false alarms" for human reviewers to sift through.

Experimental Results Table 3: Accuracy comparison showing CNN's consistent growth as training size increases.

Performance Visualization

The Precision-Recall curves below illustrate the stability of the CNN. Even as the "Recall" (the percentage of relevant documents found) increases, the CNN maintains a higher "Precision" (the accuracy of those findings) compared to the SVM baseline.

Precision-Recall Curves Figure 2: Comparative PR curves. Note the gap widening in favor of CNN in larger training sets.

Critical Insight & Practical Hurdles

While the technical superiority of CNNs for legal text is established here, the paper highlights two critical "Real World" challenges:

  1. Compute Infrastructure: SVMs can be trained on CPUs in minutes. CNNs, while more accurate, require GPU acceleration to meet the interactive speeds expected by attorneys.
  2. The Truncation Problem: CNNs generally require fixed-length inputs (1500 words in this study). Since legal documents can be massive, "chopping" the text risks losing relevant data. This points toward a need for more advanced "Long-Context" architectures in future legal AI research.

Conclusion

This study serves as a pivotal bridge between "Old School" linear legal tech and the "New School" of Deep Learning. It proves that the inductive bias of CNNs—their ability to recognize patterns in sequences—is highly compatible with the way legal relevance is determined. For firms handling massive litigations, the investment in Deep Learning specialized hardware may soon be mandatory for maintaining a competitive edge.

Find Similar Papers

Try Our Examples

  • Search for recent studies that compare Transformer-based models like BERT or Legal-BERT against CNNs for technology-assisted review (TAR) in the legal domain.
  • Which paper first introduced the benchmark comparison between deep learning and traditional SVMs for legal predictive coding, and how have those findings evolved with larger datasets?
  • How have attention mechanisms or Hierarchical Attention Networks (HAN) been applied specifically to long-form legal document classification to address the text truncation issues mentioned in this paper?
Contents
CNN vs. SVM: Scaling Predictive Coding in the Legal Frontier
1. TL;DR
2. Contextual Motivation
3. The Architecture of Legal Insight
3.1. Key Technical Components:
4. Experimental Battleground
4.1. Key Findings:
5. Performance Visualization
6. Critical Insight & Practical Hurdles
7. Conclusion