Crowd-Type: Bridging the Semantic Gap in Knowledge Bases with Human Intelligence

Crowd-Type: A Crowdsourcing-Based Tool for Type Completion in Knowledge Bases

2018-01-01
Zhaoan Dong, Jianhong Tu, Ju Fan, Jiaheng Lu, Xiaoyong Du, Tok Wang Ling
Summary
Problem
Method
Results
Takeaways
Abstract

Crowd-Type is a hybrid crowdsourcing-based system designed for fine-grained entity type completion in Knowledge Bases (KBs). It integrates automatic algorithms like SDType with human intelligence to identify and verify missing hierarchical types, significantly outperforming purely machine-based approaches in accuracy for specific, low-level categories.

TL;DR

Crowd-Type is a sophisticated system designed to solve the "fine-grained type missing" problem in Knowledge Bases (KBs). By combining the scalability of automatic algorithms with the precision of human workers, it identifies the most "influential" entities for crowdsourcing and propagates those high-quality labels across the KB using an embedding-based influence model.

Context & Positioning

In the world of Semantic Web and Linked Data (like DBpedia), broad types are common—we know a "Person" is a "Person." However, the utility of a KB depends on its specificity. Knowing that a person is a SoccerPlayer or a Physicist is what enables complex queries. Standard SOTA algorithms like SDType rely on statistical distributions of predicates, which often overlap for specific sub-types, leading to low confidence scores. Crowd-Type enters the scene not as a replacement for SDType, but as a "human-powered booster" that targets the machine's blind spots.

The Problem: The Precision Bottleneck

Automatic type inference faces a hard ceiling. When two types (e.g., "Actor" and "Director") share almost identical predicate patterns in a noisy dataset, a machine becomes "confused"—its confidence scores stagnate.

  1. Data Sparsity: Fine-grained types have fewer instances to learn from.
  2. Cost of Crowdsourcing: We cannot ask humans to verify millions of entities.
  3. Dependency: Entity types are not independent; they exist in a hierarchical ontology.

Methodology: High-Logic Hybridization

1. The Candidate Type Graph (CTG)

The system first runs an automatic predictor (SDType) to generate potential entity-type pairs. This creates a graph of possibilities where participants are represented by potential labels and confidence scores.

2. Representative Entity Selection

This is the "brain" of the system. Instead of random sampling, it selects entities based on:

  • Uncertainty: Where the machine is most unsure.
  • Influence: Which entity, if labeled, would provide the most information about its neighbors in the embedding space?

System Architecture Figure 1: The Crowd-Type Architecture, showing the loop between Task Selection and Type Inference.

3. Embedding-Based Influence Model

Once a human verifies an entity, those results aren't just applied to that one entity. The system uses an Embedding-Based Influence Model.

  • Logic: If Entity A and Entity B are close in the latent embedding space (meaning they share semantic contexts), and a human confirms Entity A is a "SoccerPlayer," the probability that Entity B is also a "SoccerPlayer" increases.

Candidate Types Graph Figure 2: The propagation mechanism where human labels (green) influence unverified candidates (blue).

Demonstration & Results

In a real-world scenario using DBpedia 3.8 (with over 2 million entities), the system demonstrates a dynamic update of the KB.

  • Worker Interface: Humans are presented with micro-tasks that include short descriptions and class hierarchies to ensure accuracy.
  • Performance Swing: The "Task State" view shows real-time propagation. When a human confirms a type, the confidence scores of "neighboring" entities in the graph receive a "⇑" (increase) or "⇓" (decrease). This illustrates that human intelligence effectively "unclogs" the machine's decision-making process.

Verification Interface Figure 3: Detailed view of how worker input shifts confidence scores for specific entities.

Critical Insight & Conclusion

The real value of Crowd-Type is its Inductive Bias management. It acknowledges that machines are great at processing the "head" of the data distribution (general types) while humans are required for the "tail" (fine-grained types).

Takeaway: This work proves that the future of Knowledge Engineering isn't about perfect algorithms, but about perfect orchestration between machine processing and human verification.

Limitations: The system heavily relies on the quality of the initial embedding space. If the embedding fails to capture the nuances of rare types, the "influence" propagation might lead to error cascading. Future iterations might benefit from integrating Large Language Models (LLMs) to provide the initial "Machine" prediction, potentially reducing the human load even further.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Active Learning and Crowdsourcing specifically for Knowledge Graph completion or refinement.
  • What is the theoretical origin of the Embedding-Based Influence Model used in Crowd-Type, and how does it compare to Graph Convolutional Networks (GCNs) for label propagation?
  • Explore how hybrid machine-crowdsourcing frameworks have been applied to multi-modal knowledge bases containing both text and images.
Contents
Crowd-Type: Bridging the Semantic Gap in Knowledge Bases with Human Intelligence
1. TL;DR
2. Context & Positioning
3. The Problem: The Precision Bottleneck
4. Methodology: High-Logic Hybridization
4.1. 1. The Candidate Type Graph (CTG)
4.2. 2. Representative Entity Selection
4.3. 3. Embedding-Based Influence Model
5. Demonstration & Results
6. Critical Insight & Conclusion