Beyond the Plot: Decoding the Human Perception of Narrative Similarity

Using Crowdsourcing to Investigate Perception of Narrative Similarity

2014-11-03
Dong Nguyen, Dolf Trieschnigg, Mariët Theune
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates how humans perceive narrative similarity in the context of Dutch folk narratives. Using crowdsourcing, the authors compare judgements from non-experts and experts, introducing a multi-dimensional analysis of story similarity and evaluating various NLP-based similarity models.

TL;DR

Information retrieval often treats document similarity as a flat vector of word counts. This paper proves that for stories, similarity is far more complex. By crowdsourcing thousands of judgements on Dutch folktales, the authors reveal that while experts look at deep structural "motifs," regular readers are influenced by genre, writing style, and the level of detail. The study shows that existing scholarly classifications (story types) only tell half the story.

The "Story Type" Problem: A Rigorous Critique

In folklore studies, stories are traditionally grouped by Story Types (e.g., "Little Red Riding Hood"). This is a binary system: Two stories either share a type or they don't.

The authors argue this is insufficient for modern digital libraries. Why? Because a joke about a cat and a dark legend about a witch-cat might share a "type" but feel completely different to a reader. The research aims to shift from these categorical labels to a continuous, multi-dimensional understanding of similarity that reflects actual human intuition.

Methodology: The Expertise Gap

The researchers conducted a large-scale empirical study using the Dutch Folktale Database. They compared two distinct groups:

  1. The Crowd: 80 workers providing ratings and motivations.
  2. The Experts: 3 senior folktale researchers.

They sampled narrative pairs across different genres (Fairy tales, Jokes, Urban legends) and lexical similarity levels (Cosine similarity bins) to ensure a diverse dataset.

Model Task Design The study utilized a specific HIT (Human Intelligence Task) design to capture both quantitative scores and qualitative motivations.

Dimensions of Similarity: What Actually Matters?

By analyzing the free-text motivations provided by participants, the authors identified several "Narrative Dimensions."

DimensionNon-Expert FocusExpert Focus
PlotHighHigh
CharactersModerateHigh
Genre (e.g., Joke vs Legend)HighLow
Style/PunctuationModerateLow
Motifs/Story TypesNoneHigh

The OLS regression analysis revealed a surprising insight: for the crowd, "Style differences" and "Difference in number of details" were statistically significant indicators of similarity ratings. If one story was a brief summary and the other a detailed narrative, users rated them as less similar—even if the plot was identical.

Can Machines Learn Story Intuition?

The paper evaluated supervised learning models to predict these human ratings.

  • Lexical Features: Jaccard index on character n-grams outperformed standard word-based Cosine similarity.
  • Story Elements: Approximating plot via Subject-Verb pairs and themes via LDA provided moderate correlations but were often "swallowed" by the high performance of lexical n-grams.
  • The Hybrid Winner: The best performance (Spearman ρ = 0.592) was achieved by combining lexical data with manually annotated metadata (keywords and story types).

Experimental Results Table 14: Performance comparison of various feature sets in predicting human similarity scores.

Critical Insight & Conclusion

The most striking takeaway is the Expert-Layman Divergence. Experts are "trained" to ignore the surface-level noise of a story (like whether it's funny or sad) to reach the structural core. Average users, however, are deeply affected by the texture of the narrative.

For developers building recommendation engines for fiction or news, the lesson is clear: Plot is not enough. To truly satisfy a user's sense of "something similar," you must account for the genre's "vibe" and the narrator's specific stylistic choices.

Limitations

The study is focused on the Dutch language and folk-specific genres. Furthermore, extracting "Plot" automatically remains a challenge—the Frog parser used in the study occasionally struggled with the non-standard language found in old legends, limiting the effectiveness of structural features.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize crowdsourcing to build gold-standard datasets for semantic text similarity beyond simple topical overlap.
  • What are the foundational theories regarding "story types" and "motifs" in computational folkloristics, and how have they been modernized in the era of LLMs?
  • Explore research that applies multi-dimensional narrative similarity models to modern content recommendation systems, such as movie scripts or news storytelling.
Contents
Beyond the Plot: Decoding the Human Perception of Narrative Similarity
1. TL;DR
2. The "Story Type" Problem: A Rigorous Critique
3. Methodology: The Expertise Gap
4. Dimensions of Similarity: What Actually Matters?
5. Can Machines Learn Story Intuition?
6. Critical Insight & Conclusion
6.1. Limitations