Costume Core: Crowdsourcing the Digital Heritage of Chinese Qipao
Crowdsourcing Metadata Schema Generation for Chinese-Style Costume Digital Library
The paper introduces a crowdsourcing-based methodology to generate specialized metadata schemas for cultural heritage, specifically for Chinese-style costumes. It presents "Costume Core" (CC), a new schema developed for the Qipao Virtual Museum and Digital Library (QiVMDL2), which outperforms general standards in descriptive depth.
TL;DR
To solve the limitations of generic metadata standards in describing specialized cultural heritage, this paper introduces a crowdsourcing methodology to build the "Costume Core" (CC) schema. By engaging a community of 100 users to rank descriptive attributes, the researchers moved beyond the rigid constraints of Dublin Core to create a search-optimized digital library for the Chinese Qipao.
The "One-Size-Fits-None" Problem in Digital Libraries
Digital Libraries (DLs) serve as vital repositories for cultural preservation. However, when it comes to the intricate world of fashion—specifically the Chinese Qipao—standard schemas like Dublin Core (too simple) and VRA Core (focused on visual arts) fall short.
If a researcher wants to search for a Qipao based on its collar style, sleeve length, or closure type, standard systems often fail because these specific attributes aren't indexed. Conventional schema generation relies on a few experts, which is slow and often biased. The authors identify a critical gap: How do we build a schema that is both professional enough for historians and intuitive enough for the general public?
Methodology: The 5-Step Crowdsourcing Pipeline
The authors propose a systematic workflow to democratize the creation of metadata.
1. Building the Metadata Pool
They didn't start from scratch. First, they harvested elements from Dublin Core and VRA Core, then added "Qipao attributes" derived from literature (e.g., Buttons, Pleats, Cutting). This created an exhaustive—but potentially redundant—list.
2. Community Voting
Instead of guessing what's important, they surveyed 100 participants (mostly students and scholars from Singapore and China). Users rated the importance of each element from 1 to 5.
3. Data-Driven Refinement
By analyzing the mean scores, the team identified which elements "actually mattered" to users. Elements with a mean score > 4.0 were retained. Interestingly, many "standard" metadata fields were discarded in favor of visual-physical attributes.
Table: Ranking of importance. Notice how "Decorative Design" and "Collar" rank higher than many administrative tags.
Architecture and Application
The final Costume Core (CC) schema consists of 14 elements divided into three categories:
- From DC: Title, Subject, Description, File-Type.
- From VRA Core: Material, Time Span, File-Format, Publication Date.
- From Qipao Attributes: Color, Decorative Design, Collar, Sleeve, Closure, Cutting.
These were implemented using the Greenstone Digital Library Software, enabling a rich, multi-faceted search experience for the QiVMDL2 (Qipao Virtual Museum and Digital Library).
Display of the digital library utilizing the new Costume Core schema.
Critical Insight: Why Crowdsourcing Works Here
The brilliance of this approach lies in its balance of Inductive Bias. While Dublin Core provides the "Universal" structure, the "Crowd" provides the "Niche" expertise.
The study found that 7 out of the top 12 most important elements weren't in the original Dublin Core. This proves that for specialized domains like ethnic costumes, experts and standards often overlook the very features users use to differentiate artefacts.
Conclusion & Future Outlook
This paper provides a blueprint for any digital archivist working with non-traditional collections. By using the "wisdom of the crowd," we can generate schemas that are:
- Objective: Reflecting the needs of the actual user base.
- Efficient: Stripping away technical "metadata bloat" that users ignore.
- Scalable: Applicable to other costume types (e.g., Kimono, Hanbok).
Future Work: While this study focused on manual voting, the next frontier is Automated Crowdsourcing—where AI models learn from user search behavior to suggest new metadata tags automatically.
Index: Metadata Schema, Qipao, Digital Library, Crowdsourcing, Cultural Heritage.
