From Paper Cards to Graph Databases: Digitizing the WWII Japanese-American Incarceration Records
Digital Curation of a World War II Japanese-American Incarceration Camp Collection: Implications for Sociotechnical Archival Systems
This paper introduces a Computational Archival Science (CAS) framework applied to the digital curation of 25,000 WWII Japanese-American Incarceration Camp index cards. It utilizes a custom NLP/NER pipeline within the DRAS-TIC repository to automate item-level metadata extraction, social network analysis, and privacy-preserving redaction.
TL;DR
This research presents a transformative approach to archival science by applying Computational Archival Science (CAS) to a massive collection of WWII Japanese-American Incarceration Camp records. By moving beyond simple scanning, the team leverages NLP pipelines, social network analysis, and scalable repositories to turn 25,000 index cards into a searchable, relational, and ethically curated digital ecosystem.
The Archive Crisis: Scale and Silence
For decades, hundreds of millions of item-level records—from WWI award cards to WWII internment rosters—have remained locked in administrative limbo. The problem isn't just that they are analog; it’s that traditional cataloging methods cannot handle the volume.
The War Relocation Authority (WRA) records are particularly sensitive. They contain "internal security" reports on Japanese-Americans imprisoned during WWII. These aren't just names; they are stories of "disorderly conduct, theft, and accidents" recorded by camp staff. Simply digitizing them as images leaves the real history hidden.
The Methodology: Archiving as a Sociotechnical System
The authors argue that archiving in the digital age is not just a technical challenge but a sociotechnical one. Their workflow doesn't just rely on code; it relies on the interaction between algorithms and human expertise.
1. The NLP/NER Pipeline
Using the GATE (General Architecture for Text Engineering) software and the ANNIE plugin, the team developed a workflow to extract more than just names. They target:
- Case Report IDs
- Housing Identifiers
- Offense Types
- Organizations and Locations

2. DRAS-TIC: Scaling the Repository
To manage this metadata, they utilized DRAS-TIC (Digital Repository At Scale That Invites Computation). Unlike traditional static repositories, DRAS-TIC uses REST APIs and NoSQL back-ends to allow near real-time searching of extracted NLP terms, effectively turning a static image into a queryable data point.
Unlocking Social Data
One of the most profound insights of this research is the use of Social Network Analysis (SNA). By linking cards that share names, dates, and incidents, the researchers can visualize the hidden social structures within the camps.

This graph-based modeling allows researchers to see how individuals were interconnected across different incident contexts, providing a much richer historical narrative than a simple list of names.
The Privacy Challenge: Protecting Internment History
Digitization often clashes with privacy. The National Archives (NARA) mandates that information regarding internees who were 18 or younger at the time of the incident must not be released.
To solve this, the researchers built Name Gazetteers using historical data files to determine the ages of individuals on the cards automatically. This "Privacy Mining" ensures that the digital release of the archive respects both historical transparency and individual ethics.
Critical Analysis & Future Outlook
The value of this work lies in its holistic framework. It acknowledges that:
- Algorithms have biases: Government-produced content is inherently biased; community involvement (e.g., Densho.org) is necessary to correct the narrative.
- Archival Operations must be preserved: The process of how an OCR error was corrected or how a metadata rule was applied is itself a historical record that belongs in an Archival Information Package (AIP).
Limitation: The current system relies heavily on the quality of OCR. While crowdsourcing helps, the sheer diversity of card styles across different camps remains a significant hurdle for universal automation.
Takeaway: This project serves as a blueprint for the future of "Big Data" in the humanities. It proves that by treating archives as dynamic sociotechnical systems, we can finally unlock the hundreds of millions of records currently gathering dust in national repositories.
