Beyond Flat Data: Uncovering Hidden Discrimination via Semantic Ontologies

Classification Rule Mining Supported by Ontology for Discrimination Discovery

2016-12-01
Binh Thanh Luong, Salvatore Ruggieri, Franco Turini
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel framework for discrimination discovery that integrates ontology engineering with generalized classification rule mining. It leverages the legal methodology of "situation testing" within a semantic web context to identify unequal treatment in historical datasets, specifically demonstrated on the U.S. Harmonized Tariff Schedules (HTS) to uncover gender-based tax disparities.

TL;DR

Researchers from the University of Pisa have developed a framework that combines Ontology Engineering with Classification Rule Mining to detect systemic bias. By applying the legal principle of "situation testing" to a semantic hierarchy, they've demonstrated how to uncover complex discriminatory patterns in the U.S. Tariff system that traditional "flat" data mining methods would likely miss.

Problem & Motivation: The Limits of Flat Analysis

Most automated discrimination discovery tools treat datasets as simple tables. However, the real world is hierarchical. For example, in the U.S. Harmonized Tariff Schedules (HTS), a "tunic" is a type of "shirt," which is a type of "apparel."

Prior works often overlooked these relationships, leading to two major issues:

  1. Lack of Context: They couldn't distinguish between legally relevant attributes (PND) and protected characteristics (PD) within a nested hierarchy.
  2. Pattern Explosion: Without semantic grouping, analysts are buried under thousands of redundant rules.

The authors' insight was to use Ontologies as the backbone. This allows the system to compare "similar" individuals—not just those with identical values, but those who sit close to each other in a conceptual tree.

Methodology: Situation Testing at Scale

The core of the paper is the translation of the legal Situation Testing (or auditing) methodology into a computational logic.

1. Semantic Similarity

Instead of simple equality, the authors define similarity through Path Distance in the ontology. If two items belong to the same sub-class, their similarity is 1. If they share a parent, it decreases exponentially ().

2. The Discriminatory Indicator

For a given context (e.g., "Outerwear made of synthetic fiber"), the system looks for pairs of realizations where:

  • The non-sensitive attributes are nearly identical (High Similarity).
  • The protected attribute (e.g., Gender) is different.
  • The outcome (e.g., Tariff rate) is significantly worse for the protected group.

3. Rule Extraction

The framework is implemented as a Protégé plugin, using SWRL (Semantic Web Rule Language) to query the knowledge base and extract generalized rules.

HTS Overall Architecture Fig 1: The proposed framework architecture, showing the flow from ontology population to rule generation.

Experiments: The HTS Case Study

The researchers applied this to the U.S. HTS dataset, specifically looking for gender bias in clothing taxes.

Key Findings:

  • Footwear & Sleepwear: These categories showed the "steepest" distributions of discriminatory rules, indicating very specific and precise contexts where women's products are taxed higher than men's.
  • Granularity Matters: The system found that "Shorts" had a higher confidence level for discrimination than the general "Pants" category. This "Multi-level" discovery allows analysts to pinpoint exactly where the bias originates.

Experimental Results Comparison Fig 2: Cumulative distribution of discriminatory rules. Note how different categories (Trousers vs. Pants) exhibit distinct bias patterns.

Critical Analysis & Conclusion

Takeaway

This work represents a significant step toward Semantic Data Mining. By using an ontology, the discovery process becomes more "human-centric" and aligned with legal reasoning. It successfully proves that gender-based tariffs in the U.S. aren't just random outliers but are embedded in specific material and form-based categories.

Limitations

  • Scalability: While performance on the HTS dataset was fast, using complex SWRL queries on massive datasets (e.g., social media logs) might hit reasoning performance bottlenecks.
  • Ontology Quality: The system is highly dependent on the "Domain Expert" to build a correct TBox. If the ontology is biased or incomplete, the discovery will be too.

Future Outlook

The authors suggest moving toward Dynamically Inferred Properties. Imagine a system that doesn't just look at static tags but uses reasoning to "understand" if a product is becoming more similar to another over time, adjusting its discrimination alerts in real-time.

Find Similar Papers

Try Our Examples

  • Search for recent papers that combine Knowledge Graphs or Ontologies with Algorithmic Fairness to detect bias in financial or legal sectors.
  • Which original studies established the "Situation Testing" methodology in data mining, and how does this paper's semantic similarity approach differ from early k-NN implementations?
  • Explore how SWRL (Semantic Web Rule Language) is being used in modern Explainable AI (XAI) to provide human-readable justifications for model decisions.
Contents
Beyond Flat Data: Uncovering Hidden Discrimination via Semantic Ontologies
1. TL;DR
2. Problem & Motivation: The Limits of Flat Analysis
3. Methodology: Situation Testing at Scale
3.1. 1. Semantic Similarity
3.2. 2. The Discriminatory Indicator
3.3. 3. Rule Extraction
4. Experiments: The HTS Case Study
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook