Leveraging Crowdsourcing Networks: A Structural Approach to Schema Matching
On Leveraging Crowdsourcing Techniques for Schema Matching Networks
This paper introduces a crowdsourcing framework for validating schema matching results in large-scale "schema matching networks." By leveraging network-level constraints (1-1 and circle constraints) and contextual question design, the system significantly enhances matching accuracy and reduces human effort.
TL;DR
Schema matching—the process of identifying correspondences between different database schemas—is a classic bottleneck in data integration. This paper shifts the focus from simple pair-wise matching to Schema Matching Networks. By treating schemas as nodes in a graph and using the crowd to validate links, the authors demonstrate that systemic constraints (like transitivity and 1-1 mappings) can cut human validation costs by 50% while ensuring high precision.
Contextual Intelligence: Beyond Pair-wise Comparison
The fundamental intuition of this work is that matching Schema A to Schema B is easier if you also know how they both relate to Schema C.
Existing tools like COMA or AMC often fail to capture the semantic nuances of attributes (e.g., is BirthName the same as Name or BirthDate?). The authors argue that by presenting "Contextual Information" to crowd workers—specifically Transitive Closures (positive evidence) and Transitive Violations (negative evidence)—the inherent ambiguity of data attributes can be resolved much faster than looking at two columns in isolation.
Methodology: The Power of Constraints
The core innovation lies in how worker answers are aggregated. Instead of simple majority voting, the paper uses an Expectation-Maximization (EM) approach that factors in worker reliability.
More importantly, it introduces justified aggregation through two primary constraints:
- 1-1 Constraint: If attribute is already matched to with high confidence, the probability of matching should decrease.
- Circle Constraint: If and are true, then is highly likely to be true (Interoperability).
System Architecture
The framework operates in a loop: generating candidates, building contextual questions, and aggregating results until an error threshold is met.

Experiments and Results
The authors validated their approach on real-world datasets including Google Fusion Tables and WebForms.
1. Contextual Impact
The study showed that providing Transitive Closure contexts helped workers confirm correct matches more easily, while Transitive Violation contexts were highly effective at helping workers reject incorrect heuristic candidates.
2. Efficiency Gains
The most striking result is the relationship between the error threshold and the cost (number of questions). By applying constraints, the system reaches the desired confidence level with significantly fewer human interventions.
Figure: The expected cost of aggregation with constraints is approximately half that of the case without constraints across various worker reliability levels.
Critical Insight & Conclusion
This paper proves that structure is a feature. In a world of fragmented data, we shouldn't just ask the crowd "is X equal to Y?". Instead, we should ask them to validate a web of logic. By converting global network consistency into local probabilistic updates, we can solve the "uncertainty problem" of automated matchers without breaking the bank.
Limitations: The current model focuses on 1-1 and circle constraints. Future work could benefit from exploring more complex functional dependencies or domain-specific logic (e.g., geographical constraints) to further prune the search space.
