Beyond the Plot: Decoding the Human Perception of Narrative Similarity
Using Crowdsourcing to Investigate Perception of Narrative Similarity
This paper investigates how humans perceive narrative similarity in the context of Dutch folk narratives. Using crowdsourcing, the authors compare judgements from non-experts and experts, introducing a multi-dimensional analysis of story similarity and evaluating various NLP-based similarity models.
TL;DR
Information retrieval often treats document similarity as a flat vector of word counts. This paper proves that for stories, similarity is far more complex. By crowdsourcing thousands of judgements on Dutch folktales, the authors reveal that while experts look at deep structural "motifs," regular readers are influenced by genre, writing style, and the level of detail. The study shows that existing scholarly classifications (story types) only tell half the story.
The "Story Type" Problem: A Rigorous Critique
In folklore studies, stories are traditionally grouped by Story Types (e.g., "Little Red Riding Hood"). This is a binary system: Two stories either share a type or they don't.
The authors argue this is insufficient for modern digital libraries. Why? Because a joke about a cat and a dark legend about a witch-cat might share a "type" but feel completely different to a reader. The research aims to shift from these categorical labels to a continuous, multi-dimensional understanding of similarity that reflects actual human intuition.
Methodology: The Expertise Gap
The researchers conducted a large-scale empirical study using the Dutch Folktale Database. They compared two distinct groups:
- The Crowd: 80 workers providing ratings and motivations.
- The Experts: 3 senior folktale researchers.
They sampled narrative pairs across different genres (Fairy tales, Jokes, Urban legends) and lexical similarity levels (Cosine similarity bins) to ensure a diverse dataset.
The study utilized a specific HIT (Human Intelligence Task) design to capture both quantitative scores and qualitative motivations.
Dimensions of Similarity: What Actually Matters?
By analyzing the free-text motivations provided by participants, the authors identified several "Narrative Dimensions."
| Dimension | Non-Expert Focus | Expert Focus |
|---|---|---|
| Plot | High | High |
| Characters | Moderate | High |
| Genre (e.g., Joke vs Legend) | High | Low |
| Style/Punctuation | Moderate | Low |
| Motifs/Story Types | None | High |
The OLS regression analysis revealed a surprising insight: for the crowd, "Style differences" and "Difference in number of details" were statistically significant indicators of similarity ratings. If one story was a brief summary and the other a detailed narrative, users rated them as less similar—even if the plot was identical.
Can Machines Learn Story Intuition?
The paper evaluated supervised learning models to predict these human ratings.
- Lexical Features: Jaccard index on character n-grams outperformed standard word-based Cosine similarity.
- Story Elements: Approximating plot via Subject-Verb pairs and themes via LDA provided moderate correlations but were often "swallowed" by the high performance of lexical n-grams.
- The Hybrid Winner: The best performance (Spearman ρ = 0.592) was achieved by combining lexical data with manually annotated metadata (keywords and story types).
Table 14: Performance comparison of various feature sets in predicting human similarity scores.
Critical Insight & Conclusion
The most striking takeaway is the Expert-Layman Divergence. Experts are "trained" to ignore the surface-level noise of a story (like whether it's funny or sad) to reach the structural core. Average users, however, are deeply affected by the texture of the narrative.
For developers building recommendation engines for fiction or news, the lesson is clear: Plot is not enough. To truly satisfy a user's sense of "something similar," you must account for the genre's "vibe" and the narrator's specific stylistic choices.
Limitations
The study is focused on the Dutch language and folk-specific genres. Furthermore, extracting "Plot" automatically remains a challenge—the Frog parser used in the study occasionally struggled with the non-standard language found in old legends, limiting the effectiveness of structural features.
