Breaking the Silence on Label Noise: How Clustering and Self-Training Outperform Traditional Data Polishing

Expert Systems With Applications

2025-01-01
Som Gupta
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces two novel algorithms for correcting label noise: Self-Training Correction (STC) and Cluster-based Correction (CC). Evaluated against the adapted "Polishing Labels" (PL) method, these algorithms demonstrate significant improvements in label accuracy, model quality, and AUC across various datasets, particularly in crowdsourcing environments.

TL;DR

While data filtering has long been the "go-to" for noisy datasets, it often throws the baby out with the bathwater. This paper introduces Cluster-based Correction (CC) and Self-Training Correction (STC)—two proactive methods that fix incorrect labels instead of deleting them. Experiments show these methods, especially CC, significantly boost model performance in the messy world of crowdsourced data.

Problem & Motivation: The Over-Cleansing Trap

In the era of crowdsourcing (e.g., Amazon Mechanical Turk), we rely on non-experts who inevitably make mistakes. The standard solution has been Noise Filtering—detecting and deleting suspicious rows. However, this leads to over-cleansing, where the resulting dataset is too small to train a robust model.

The authors identify a gap: Label Correction. Instead of discarding data, can we systematically flip wrong labels back to their ground truth? Previous attempts like "Polishing" were computationally heavy and modified feature attributes, which can corrupt the underlying data distribution. The goal here is a "Label-Only" correction that preserves the feature space.

Methodology: The Core Strategies

The paper contrasts three specific correction philosophies:

1. Polishing Labels (PL)

A simplified version of Teng’s 1999 algorithm. It uses a 10-fold cross-validation approach where an ensemble of classifiers "votes" on what the label should be. If the consensus disagrees with the original label, the label is changed.

2. Self-Training Correction (STC)

Borrowing from semi-supervised learning, STC treats a "cleaned" subset (identified by a filter) as labeled data and the "noisy" subset as unlabeled. It iteratively trains a classifier and "re-labels" the noisy instances for which it has the highest confidence.

3. Cluster-Based Correction (CC) - The Winner

This is the paper’s most innovative contribution. Unlike the others, CC is unsupervised. It groups data points based on feature similarity (K-means). It assumes that if a cluster is dominated by a specific class, outliers in that cluster are likely noise.

  • Weighted Approach: It doesn't just look at one cluster; it runs K-means 200 times with different seeds and "k" values.
  • Intuition: By aggregating weights across hundreds of different clusterings, the true signal emerges from the noise.

Cluster-based Correction Pseudo-code

Experimental Results: CC Takes the Crown

The authors tested these methods on both artificial noise (simulated) and real-world crowdsourcing data (Leaf images).

  • Label Quality: CC consistently improved accuracy across both binary and multi-class datasets.
  • Model Quality: Interestingly, the study found that while traditional Polishing (PL) sometimes harmed model quality, CC almost always improved it.
  • Crowdsourcing Synergy: When combined with consensus methods like Majority Voting (MV) or Dawid-Skene (DS), CC acted as a powerful refinement step, often making simple voting as effective as complex statistical models.

Average Accuracy Comparison Figure: CC (Blue) consistently stays above the baseline (Dashed line) more effectively than PL (Green) or STC (Red).

Critical Analysis & Conclusion

The standout takeaway is the robustness of unsupervised clustering. While classifiers (STC, PL) are prone to inheriting the bias of the noisy labels they are trained on, clustering looks at the inherent "spatial" structure of the data, making it less susceptible to being misled by wrong labels.

Limitations & Future Work

  • Low Noise Threshold: Correcting labels when noise is very low (below 10%) remains a challenge; correction can accidentally "corrupt" clean labels.
  • Computational Cost: Running K-means hundreds of times is effective but might be slow for billion-scale datasets.
  • Future Path: The authors suggest automating the selection of parameters (like the number of clusters) to make this a "hands-off" pre-processing tool for data scientists.

Bottom line: If you are working with crowdsourced labels, don't just filter—correct with clustering. It’s a more efficient way to squeeze every bit of value out of your expensive data.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Cluster-based Correction (CC) using deep clustering or ensemble manifold learning to handle non-spherical high-dimensional label noise.
  • Which study first introduced the "Polishing" concept for data cleaning, and how does this paper's "Polishing Labels" simplification affect the trade-off between attribute integrity and classification accuracy?
  • Explore how the proposed label noise correction methods have been applied to modern deep learning pipelines for cleaning massive-scale web-crawled datasets or medical imaging annotations.
Contents
Breaking the Silence on Label Noise: How Clustering and Self-Training Outperform Traditional Data Polishing
1. TL;DR
2. Problem & Motivation: The Over-Cleansing Trap
3. Methodology: The Core Strategies
3.1. 1. Polishing Labels (PL)
3.2. 2. Self-Training Correction (STC)
3.3. 3. Cluster-Based Correction (CC) - *The Winner*
4. Experimental Results: CC Takes the Crown
5. Critical Analysis & Conclusion
5.1. Limitations & Future Work