From Archival Strings to Tensors: Deep Learning the Victorian Entrepreneur
Information Processing and Management
This paper presents a robust binary classification framework for identifying entrepreneurs in British historical census data (1851–1881) using the I-CeM dataset. By leveraging advanced Machine Learning and Deep Learning (LSTM-RNN) architectures, the researchers achieved a near-perfect classification accuracy of 0.9978, significantly outperforming traditional econometric benchmarks.
TL;DR
History meets High-Tech. This paper demonstrates how to recover lost economic status for 80 million people from 19th-century British censuses. By replacing traditional Logistic Regression with Deep Learning (LSTM-RNN), the authors pushed classification accuracy from a mediocre 74% to a near-perfect 99.78%, effectively "filling in" missing historical records with modern precision.
Problem & Motivation: The Silent Entrepreneurs
In the mid-Victorian era (1851–1881), census takers didn't ask "Are you an employer?" They simply asked for an occupation. An individual might write "Boot and shoe maker employing 4 men" or just "Master Carpenter."
For decades, economic historians used Logistic Regression (LR) to guess employment status based on age or marital status. However, LR is a "linear" tool trying to solve a "non-linear" world. It fails to capture the "florid" nuances of the 500-character occupational strings (OccString), leading to significant misclassification and skewed historical narratives about entrepreneurship.
Methodology: The "Uncrumpling" of Data
The authors propose a transition from shallow to deep learning. The core insight is that occupational data behaves like a folded manifold.
1. The Classical Baseline: Logistic Regression
Used as the "Hello World" of their experiment, the baseline reached 74% accuracy. It treated variables like "Age" and "Number of Servants" as independent factors, missing the complex interactions between them.
2. The Ensemble Leap: AdaBoost
By using AdaBoost, the researchers built a "committee" of weak learners. This model significantly outperformed LR by focusing on hard-to-classify cases, reaching 95% accuracy.
3. The Core: RNN with LSTM
To truly master the OccString, the authors deployed a Recurrent Neural Network (RNN) with an LSTM (Long Short-Term Memory) layer. Unlike the "Bag of Words" approach (which treats words like an unordered soup), the LSTM captures the sequence and context of the professional description.
Fig 1: The jump from simple Logistic Regression (LR) to AdaBoost shows the significant reduction in False Positives.
Experiments & Results: Slaying the Baseline
The study compared ten optimized algorithms. The results were clear: Standard algorithms are systematically outperformed by ML/DL.
- Logistic Regression: 0.74 Accuracy (The floor)
- AdaBoost: 0.95 Accuracy
- Deep Learning (Dense): 0.96 Accuracy
- LSTM-RNN (Text-aware): 0.9978 Accuracy (The ceiling)
Fig 2: 2-D predicted probability grids show how different models (like SVM or Random Forest) "carve" the decision space for entrepreneurs (Green) vs. Workers (Purple).
The ROC Curves (Receiver Operating Characteristic) further confirmed that Gradient Boosting and Ensemble methods achieved the "Top Left Corner"—the holy grail of classification where True Positives are maximized and False Alarms are virtually eliminated.
Critical Insight: Why Does DL Win Here?
The paper utilizes the metaphor of "uncrumpling paper." Historical data is "dirty" and "folded." Traditional regression tries to draw a straight line through a crumpled ball. Deep Learning, through its multiple hidden layers and non-linear activations (like ReLU and Sigmoid), systematically "smoothes out" the paper until the entrepreneurs and workers are clearly separated in vector space.
Conclusion & Future Outlook
This isn't just a win for history; it's a blueprint for Information Science. The success of the LSTM model suggests:
- Semantics Matter: The "OccString" was the most valuable feature.
- Beyond OLS: Social scientists must move beyond OLS and Logit if they want to handle "Big Data" censuses.
- Historical Continuity: We can now link 19th-century data to 21st-century trends with high confidence, revealing that the Victorian era was actually the "Age of Entrepreneurship" in Britain.
Final Takeaway: When historical data is messy, don't clean it—Deep Learn it.
