Subjective KBC: Bridging the Gap Between Hard Facts and Human Opinions

Subjective Knowledge Base Construction Powered By Crowdsourcing and Knowledge Base

2018-05-25
Hao Xin, Rui Meng, Lei Chen
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a two-stage framework for Subjective Knowledge Base Construction (SKBC) that integrates Probase for core structure and DBpedia for enrichment. It utilizes crowdsourcing and partial-order inference to acquire subjective facts like "(New York, big City, yes)" at scale and with high quality.

TL;DR

While modern Knowledge Bases (KBs) are masters of objective facts (e.g., "Paris is in France"), they are largely "opinion-blind." This paper presents a framework to build Subjective Knowledge Bases by combining the structured data of DBpedia/Probase with the wisdom of the crowd. By using Partial-Order Relationships, the system can infer subjective labels (like "big," "famous," or "competitive") for thousands of entities while only paying for a fraction of them to be manually annotated.

The "Opinion Gap" in Modern AI

Knowledge Graph construction has historically focused on the observable and the factual. However, a significant portion of user intent is subjective. When a user searches for a "large company," they aren't just looking for a revenue number; they are looking for entities that satisfy the human consensus of "large."

The challenge? Subjective knowledge has no ground truth in some database; the truth lives in the collective mind of the crowd. Traditional KBC methods fail here because they lack a mechanism to aggregate "dominant opinions" cost-effectively.

Methodology: High-Logic Crowdsourcing

The authors don't just ask the crowd to label everything. They use a two-step process:

1. Core Construction & Property Mapping

They mine Probase to identify "ST pairs" (Subjective Property-Type pairs) like <big, City>. They then map these to DBpedia to pull in objective features (Area, Population, etc.) that might correlate with the subjective label.

2. Inferring via Partial-Order

This is the "secret sauce." The authors define a semantic order. If we are looking for "big cities," and we know:

  • City A is "big" (annotated by crowd).
  • City B has a larger area and more people than City A.
  • Therefore, City B is also "big" (inferred).

Workflow of Subjective KB Construction

Cost-Aware Strategy: Adaptive vs. Batch

To minimize the "crowd budget," the paper treats instance selection as a mathematical optimization problem. They prove that their utility function is Adaptive Submodular, meaning we get "diminishing returns" for each new annotation.

  • Adaptive Annotation: Selects instances one by one, reacting to the last worker's answer. This is the most cost-efficient.
  • Batch-Mode: Selects groups of instances to reduce latency (waiting for workers on Amazon Mechanical Turk).

Experimental Proof

The researchers tested this on 177 gradable adjectives. The results showed that by using the Partial-Order logic, the system maintained high accuracy while drastically reducing the number of human tasks needed.

Instance Annotation Strategy Cost

As shown in the cost comparison, the Adaptive and Batch strategies significantly lowered the "HIT cost" (Human Intelligence Tasks) required to fully populate the KB compared to random sampling.

Performance vs. Machine Learning

Interestingly, the framework's inference rules outperformed pure machine learning models like SVMs and Decision Trees (83.7% vs 70.5-79.6% accuracy). This suggests that human-derived logic (partial ordering) is a stronger inductive bias for subjective tasks than standard statistical features.

Critical Insight & Future Outlook

This paper effectively turns "subjectivity" into a structured, computable domain. By linking objective metrics (area, employees, endowment) to subjective adjectives (big, large, famous), it creates a bridge between quantitative data and qualitative human language.

Limitations: The model relies on "gradable" adjectives. It might struggle with highly polarized or culturally dependent subjectivity (e.g., "beautiful") where a "dominant opinion" may not exist or may shift rapidly over time.

Takeaway: The future of KBs is not just about what is true, but what is perceived. This framework is a vital step toward search engines that truly "understand" human adjectives.

Find Similar Papers

Try Our Examples

  • Find recent papers or SOTA methods that address "subjective knowledge extraction" from unstructured web text or social media beyond crowdsourcing.
  • Which paper first introduced the concept of "Adaptive Submodularity" in the context of active learning, and how does this paper adapt that theory for cost-aware crowdsourcing?
  • Are there any studies that apply this partial-order annotation framework to multi-modal tasks, such as determining subjective attributes of images or videos?
Contents
Subjective KBC: Bridging the Gap Between Hard Facts and Human Opinions
1. TL;DR
2. The "Opinion Gap" in Modern AI
3. Methodology: High-Logic Crowdsourcing
3.1. 1. Core Construction & Property Mapping
3.2. 2. Inferring via Partial-Order
4. Cost-Aware Strategy: Adaptive vs. Batch
5. Experimental Proof
5.1. Performance vs. Machine Learning
6. Critical Insight & Future Outlook