AncestryAI: Modernizing Genealogy with Probabilistic Record Linkage

AncestryAI: A Tool for Exploring Computationally Inferred Family Trees

2017-01-01
Eric Malmi, Marko Rasa, Aristides Gionis, A. Gionis
Summary
Problem
Method
Results
Takeaways
Abstract

AncestryAI is an open-source web-based tool and framework designed to automatically reconstruct large-scale family trees from historical parish records. By employing a probabilistic record-linkage method, it transforms millions of unstructured birth and baptism records into a searchable genealogical graph.

TL;DR

AncestryAI is a sophisticated open-source platform that automates the reconstruction of family trees from massive historical datasets. By shifting from manual "search-and-match" to probabilistic inference, it enables the creation of genealogical graphs spanning millions of nodes, providing a powerful resource for both hobbyist genealogists and computational social scientists.

Problem: The Needle in the Histographic Haystack

Genealogical research has historically been a manual "detective" process. Researchers must sift through parish registers, often dealing with:

  • Spelling Noise: Variations like "Paul" vs. "Paulus" or phonetic transcriptions.
  • Demographic Overlap: Hundreds of individuals with the same common names living in the same era.
  • Fragmentation: Records are often localized, making it nearly impossible to track an ancestor who moved from one province to another without exhaustive searching.

The authors identified that while digitization projects (like Finland's HisKi) have made records accessible, they haven't made them connected.

Methodology: Bayesian Logic Meets Historical Records

The core innovation of AncestryAI lies in how it quantifies the "probability of relatedness." Instead of a binary match/no-match system, it treats record linkage as a probabilistic inference problem.

1. Probabilistic Ranking

Using a modified Fellegi–Sunter model, the tool calculates the likelihood that a specific child's birth record belongs to a set of candidate parent records. It considers three primary attributes:

  • String Similarity: Utilizing Jaro-Winkler distance to handle typos.
  • Spatial Proximity: Using coordinates of birthplaces to weight the likelihood.
  • Temporal Logic: A strict prior that parents must be between 10 and 70 years older than their children.

2. Architecture and Layout

To handle the scale of millions of records, the tool uses blocking. It clusters names and birth years so the algorithm only compares plausible candidates rather than every person in the database.

The visualization is equally innovative. Because family trees are technically Directed Acyclic Graphs (DAGs) (and not simple trees, due to intermarriage), the authors developed a heuristic layout algorithm that fixes Y-coordinates by birth year and dynamically adjusts X-coordinates to minimize edge crossings during exploration.

Model UI and Architecture Figure 1: The AncestryAI interface showing the inferred family tree (A), geographic distribution (B), and advanced search options (E).

Experiments and Results: Quantifying Family Connections

The model was validated using a "ground truth" dataset of 64,208 individuals manually verified by a professional genealogist.

The research highlights how sensitive these probabilities are to specific attributes. For example, as seen in the likelihood ratio analysis, an "identical" name match significantly boosts the matching probability, but even a slight drop in Jaro-Winkler similarity (e.g., to 0.85) serves as a strong signal against a match.

Likelihood Ratios for Name Similarity Figure 2: The weight of first-name similarity in determining match probability. Ratios above 1.0 support a match, while below 1.0 suggest the records represent different people.

Critical Insight: Beyond the Tree

The true value of AncestryAI isn't just helping someone find their great-great-grandfather. It is about Computational Social Science. By building a graph of millions of Finnish citizens over 300 years, researchers can now ask:

  • How did marriage patterns change during the industrial revolution?
  • Can we track the genetic spread of specific diseases through a computationally-verified lineage?
  • What were the actual migration patterns of the 18th-century working class?

Limitations & Future Work

Currently, the model treats each match independently. The authors acknowledge this as a limitation—in reality, a couple usually has multiple children together. Future versions will likely explore Collective Entity Resolution, where the existence of siblings reinforces the probability of the parental link, creating a more robust "family-level" inference rather than just "pair-level" matching.

Conclusion

AncestryAI bridges the gap between digital archives and actionable knowledge. It demonstrates that the same probabilistic tools we use for modern data deduplication can unlock the secrets of our biological and social history, turning fragmented church registers into a unified map of human connection.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply Collective Entity Resolution (CER) to genealogical or biographical datasets to improve link consistency across families.
  • What are the foundational principles of the Fellegi–Sunter model for record linkage, and how have modern machine learning approaches modified these priors?
  • Find studies that use computationally inferred family trees for analyzing social mobility or migration patterns over multiple centuries.
Contents
AncestryAI: Modernizing Genealogy with Probabilistic Record Linkage
1. TL;DR
2. Problem: The Needle in the Histographic Haystack
3. Methodology: Bayesian Logic Meets Historical Records
3.1. 1. Probabilistic Ranking
3.2. 2. Architecture and Layout
4. Experiments and Results: Quantifying Family Connections
5. Critical Insight: Beyond the Tree
5.1. Limitations & Future Work
6. Conclusion