Building the First Chinese Argumentation Corpus: A Crowdsourcing Triumph in Review Mining

Crowdsourcing argumentation structures in Chinese hotel reviews

2017-10-01
Mengxue Li, Shiqiang Geng, Yang Gao, Shuhua Peng, Haijing Liu, Hao Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the first Chinese argumentation corpus for hotel reviews, comprising 4,814 argument components and 411 relations. The authors propose an extended "Premise-Claim" model and utilize a crowdsourcing-based annotation workflow integrated with a K-means clustering algorithm for label aggregation.

Executive Summary

TL;DR: This paper presents the first large-scale Chinese argumentation corpus derived from hotel reviews. By leveraging crowdsourcing and an innovative K-means clustering aggregation method, the authors successfully mapped complex discourse structures—claims, premises, and relations—within the noisy, informal landscape of customer feedback.

Positioning: This work bridges the gap between expert-heavy argumentation theory and practical, large-scale data engineering. It moves beyond the traditional "legal/essay" datasets into the commercially vital but technically challenging domain of user-generated content (UGC).

Problem & Motivation: The "Chaos" of Customer Reviews

Argumentation mining (AM) is the art of teaching machines to understand why someone holds an opinion. While we have robust corpora for English persuasive essays, Chinese customer reviews present unique hurdles:

  • Informal Structure: Unlike legal briefs, reviews lack standard punctuation and consistent logical flow.
  • Implicit Logic: Users often state facts (e.g., "The subway is 5 mins away") that serve as evidence for unstated claims ("The location is good").
  • Annotation Subjectivity: What one person sees as a "Premise," another might see as a "Claim," leading to low Inter-Rater Agreement (IRA).

The authors recognized that to build a Chinese AM system, they first needed a data pipeline that could handle the inherent messiness of crowdsourced labor and subjective interpretation.

Methodology: The Extended Argumentation Model

The authors didn't just use a basic "Claim-Premise" structure. They extended it to fit the nuances of reviews:

  1. MajorClaim: Overall sentiment (e.g., "Highly recommended!").
  2. Claim: Specific attributes (e.g., "The bed was soft").
  3. Premise: Facts supporting a claim.
  4. PSIC (Premise Supporting Implicit Claim): Facts supporting a claim that isn't explicitly written.

Architecture of Annotation Aggregation

To solve the problem of conflicting human labels, the authors treated annotation as a clustering problem. Instead of simple voting, they represented each character's label as a vector and used K-means clustering to find the "centroid" of opinion. This naturally handled boundary disputes (where a sentence starts/ends) and provided a confidence score for every label.

Model Architecture and Model Types Figure 1: The proposed argumentation model depicting legal relations between MajorClaims, Claims, and Premises.

Experiments & Results: Quality Through Filtering

The study involved 388 students. To ensure quality, the authors used "Gold Standard" reviews to identify and remove "less-devoted" annotators (those who missed easy examples).

Key Findings:

  • Quality Boost: Removing low-quality workers improved the agreement score (αU) significantly across the board.
  • The "Easy vs. Controversial" Split: The authors cleverly split the results into an "Easy Reviews Corpus" (where agreement was naturally high) and a "Less-Controversial Sentences Corpus" (extracting clear nuggets from otherwise messy reviews).

Agreement Metrics Table Table 3: IRA scores confirming that for "Easy" reviews, the crowdsourced quality is comparable to expert-annotated English datasets.

Error Analysis

The most significant confusion occurred between PSIC (Implicit Premises) and Non-Argumentative (NA) text. This highlights the "interpretation gap"—some annotators see a mention of a "nearby bar" as relevant evidence for hotel quality, while others see it as irrelevant chatter.

Critical Insight & Conclusion

The true value of this paper lies in its probabilistic approach to truth. By introducing confidence scores for annotations, the authors acknowledge that human language logic isn't always binary.

Takeaway: If you are building a dataset for a subjective task, don't force a consensus where none exists. Use clustering to find the center of gravity and maintain confidence scores to help your downstream model understand which examples are "hard cases."

Limitations: The reliance on university students (novice annotators) still requires heavy post-processing. Future work could benefit from providing workers with a predefined list of "Implicit Claims" to reduce the confusion between PSIC and NA labels.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize clustering or generative AI to aggregate noisy crowdsourced natural language processing labels.
  • What are the state-of-the-art models for Chinese argumentation mining published after 2017 that utilize this specific hotel review corpus?
  • Explore research on "Premise Supporting Implicit Claim" (PSIC) detection in multi-domain customer reviews using large language models.
Contents
Building the First Chinese Argumentation Corpus: A Crowdsourcing Triumph in Review Mining
1. Executive Summary
2. Problem & Motivation: The "Chaos" of Customer Reviews
3. Methodology: The Extended Argumentation Model
3.1. Architecture of Annotation Aggregation
4. Experiments & Results: Quality Through Filtering
4.1. Error Analysis
5. Critical Insight & Conclusion