Scaling Human Intelligence for Knowledge Base Integration: The LTQS Approach
Knowledge Base Semantic Integration Using Crowdsourcing
This paper introduces a hybrid framework for Knowledge Base (KB) semantic integration using crowd intelligence. It addresses the challenges of semantic heterogeneity in taxonomies by proposing a Local Tree Based Query Selection (LTQS) mechanism and a blocking-based instance matching strategy, achieving state-of-the-art accuracy in integrating large-scale KBs like YAGO and DBpedia.
TL;DR
Integrating massive Knowledge Bases (KBs) like DBpedia and YAGO remains a challenge due to semantic mismatches. This paper presents a novel framework that leverages crowdsourcing via a Local Tree (LT) query model to resolve taxonomy hierarchies. By treating taxomony integration as a utility maximization problem, the authors achieve high-precision alignment and efficient blocking-based instance matching.
Problem & Motivation: The "Granularity Gap"
Automated tools often fail because they treat KB integration as a simple mapping problem. However, real-world KBs are heterogeneous. For instance, KB1 might categorize "Singer" under "Person," while KB2 places it under "Artist" -> "Musician."
Traditional machine-based methods achieve about 80% accuracy because they view nodes in isolation. The authors' insight is that humans need context to make accurate judgments. By showing a worker not just node A and B, but node A and the immediate children subtree of B, the system can extract much richer semantic relationships (Equivalence, Specification, or Generalization).
Methodology: Local Tree Based Query Selection (LTQS)
1. The Local Tree Query Model
Instead of asking "Is Node A the same as Node B?", the system presents a Local Tree (LTi). The crowd determines where the query node fits within that subtree.

2. Optimizing the Budget (LTQS Problem)
Since crowdsourcing costs money, the system must select the "best" questions. The authors define Utility as the Expected Pruning Power. If we confirm a relationship at a high level in the taxonomy, we can automatically prune thousands of potential child-pair comparisons (based on the Pruning Principle).
They proposed two main algorithms:
- Static Query Selection: Selects a batch of questions based on a submodular utility function ( approximation).
- Adaptive Query Selection: Re-calculates utility after every answer, allowing the system to pivot based on what the crowd has already revealed.
Experiments and Results
The researchers tested their framework on the YAGO-DBpedia alignment task.
Taxonomy Refinement
The Adaptive Selection strategy proved superior, reducing the number of candidate pairs significantly more than random selection. It successfully "fixed" class positions that were left ambiguous by automated tools like PARIS.

Instance Matching via Blocking
By using the integrated taxonomy to "block" instances (i.e., only comparing entities in "s-hop" neighboring classes), the system avoided the complexity nightmare of matching millions of records.
- Precision: Reached ~95% with a narrow block size ().
- Recall: Improved to 58% as the block size () increased, proving that a well-integrated taxonomy is the backbone of efficient data cleaning.
Depth Insight: Beyond Simple Mappings
The true value of this paper lies in its mathematical formalization of Pruning Power. By treating taxonomy as a DAG, the authors turned a qualitative human task into a quantitative optimization problem.
Limitations: While powerful, the method assumes the existence of some instance overlap to calculate "Prior Beliefs." In cold-start scenarios where KBs share no instances, the initial query selection might be less efficient.
Conclusion
The LTQS framework proves that the crowd's strength isn't just in volume, but in handling contextual hierarchy. For developers building large-scale data lakes or RAG (Retrieval-Augmented Generation) systems, this paper provides a blueprint for using human-in-the-loop systems to clean and structure messy, heterogeneous data schemas.
