Bridging Text and Expression: Building a Japanese Emotion Ontology from the Social Web
Building of Japanese Emotion Ontology from Knowledge on the Web for Realistic Interactive CG Characters
The paper introduces a novel framework for building a Japanese Emotion Ontology by mining knowledge from web sources like Twitter and BBS. It utilizes a Naïve Bayes approach to calculate emotional intensities and represents the data using OWL and EmotionML to drive realistic facial animations for CG characters.
TL;DR
In the quest for realistic interactive CG characters, the ability to "feel" and express emotions based on conversation is paramount. This paper presents a methodology to build a Japanese Emotion Ontology by mining over a million entries from Twitter and BBS (2-Channel). By calculating emotional intensities using weighted probability and structuring them via OWL and EmotionML, the authors provide a bridge that transforms raw Japanese text into nuanced 3D facial animations.
Background: The Role of Emotion in Interaction
As digital concierges and game characters become more integrated into our lives, the "Uncanny Valley" often looms large—not just because of visual fidelity, but because of emotional misalignment. Most existing systems treat emotions as binary (Happy or Sad), but human expression is a spectrum of intensity.
The authors identify a critical gap: the lack of a structured, intensity-aware emotion dataset for the Japanese language. Their solution? Leveraging the "wisdom—and emotion—of the crowds" found on social media.
Methodology: Quantifying the Ineffable
The core innovation lies in how the researchers quantified the "strength" of an emotion associated with a word.
1. The Triple-Layer Emotion Model
Instead of choosing one theory, the authors used three to ensure the ontology's versatility:
- Discrete Categories: 10 distinct emotions (Joy, Anger, Sadness, Fear, Shame, Like, Disgust, Exciting, Comforted, Surprise).
- Simple Polarity: Positive, Negative, and Neutral.
- Dimensional Model (PAD): Pleasure, Arousal, and Dominance.
2. Intensity through Probability
The authors didn't just label words; they calculated a Weighted Conditional Probability. By analyzing the occurrence frequency of words within manually tagged emotional documents, they derived a score representating how strongly a word (like Tanoshii - Joyful) correlates to a specific emotion.
Figure 1: The research pipeline from web-crawling to ontology building and application.
Architecture: From Data to Ontological Knowledge
Using OWL (Web Ontology Language), the researchers ensured that their Japanese dataset could be "linked" or merged with other global ontologies (e.g., English or Chinese datasets). Within these OWL classes, they embedded EmotionML, a W3C standard, to store the calculated intensity values as machine-readable attributes.
Calculating Variation for Animation
The bridge to the physical world (or the digital-physical world) is the MPEG-4 Specification. The system maps the highest-intensity emotion from a sentence to 68 Face Animation Parameters (FAPs).
Figure 2: Example facial animations: Neutral, Joy, Sadness, and Anger generated from ontology parameters.
Experimental Results
The authors extracted 1,033,511 texts from Twitter and 2-Channel. By focusing on adjectives—the primary carriers of emotional weight in Japanese—and applying morphological analysis (CaboCha), they created a robust intensity matrix.
A sample result for the word "Joyful" (Raku/Tanoshii) showed a high JOY intensity (0.183), while a neutral sentence like "Today was a joyful day" yielded a lower consolidated probability across multiple categories, demonstrating the system's ability to handle context.
Critical Analysis & Future Directions
Takeaway
This work is a significant step toward semantic emotional interoperability. By using OWL, the authors aren't just building a database; they are creating a component of the "Emotional Semantic Web" where different AI agents can share a common understanding of human feelings across languages.
Limitations
- Sarcasm & Context: The current Naïve Bayes approach treats words largely in isolation or simple combinations, which may struggle with the deep sarcasm common on platforms like 2-Channel.
- Morphological Focus: By focusing primarily on adjectives, the system might miss nuanced emotional cues found in Japanese verbs or sentence-ending particles.
The Path Forward
The authors propose expanding this to include time, location, and personal context, moving toward a "Linked Life Data" approach. For CG characters, the next frontier isn't just a moving face, but a "moving body" (body motion generation) that matches the ontological state of the character.
Conclusion: This paper proves that the chaotic data of the web is a goldmine for building structured knowledge that makes our digital companions more human.
