ERM4CT 2016: Bridging the Gap Between Emotion Theory and Multimodal Companion Systems
5357_ERM4CT 2016 2nd international workshop on emotion representations and modelling for companion systems (workshop summary).
This report summarizes the 2nd International Workshop on Emotion Representations and Modelling for Companion Systems (ERM4CT 2016), held at ICMI. The workshop introduces a novel 10-modality dataset designed to advance user-adaptable HCI and companion systems by focusing on multimodal affective behavior modeling.
TL;DR
The ERM4CT 2016 workshop addresses the fragmentation in Affective Computing where emotion models are often siloed by modality or discipline. By introducing the IGF-Corpus—a high-dimensional, 10-modal dataset—the organizers provide a framework for creating "Companion Systems" that possess a deep, multimodal understanding of a user’s affective state, cognitive load, and personality.
Academic Positioning: This work serves as a pivotal bridge between theoretical emotion modeling and practical implementation in user-adaptive HCI, emphasizing interoperability and naturalistic data collection.
The Interoperability Crisis in Affective Computing
Prior to the ERM4CT series, a significant bottleneck existed in HCI: emotion representations were too specific. A model designed for facial expression analysis rarely "talked" to a model designed for physiological skin conductance. This lack of a unified language makes it nearly impossible to build true Companion Systems—technologies that don't just react to commands, but adapt to a user's long-term needs and immediate affective states.
The authors identified two major hurdles:
- Modality Silos: Features used to describe emotions in audio often don't align with those used in bio-signal processing.
- The Engagement Gap: Subjets in legacy datasets were often "acting" or "passive," leading to data that lacked the nuance of real-world interactions.
Methodology: The "Gait-Training" Insight
To solve the engagement gap, the researchers embedded data collection into a functional Health and Fitness scenario.
The IGF-Corpus Architecture
Instead of asking users to "act happy," the system involved subjects (aged 50+) in a complex gait-training task. The emotional data was captured during a separate interaction with a technical interface where:
- Wizard of Oz Experiments: The system appeared autonomous (using TTS and voice commands) but was controlled to trigger specific states.
- Target States: Beyond basic emotions, the study focused on Dispositional States such as:
- Interest
- Cognitive Underload vs. Overload
- HMI-specific reactions like Frustration and Joy.

Analyzing Multi-Modal Interdependencies
The core contribution of the workshop was the "hands-on" dataset evaluation. Researchers were encouraged to look for intra- and intermodality interdependencies.
Technical Insight: If a single physiological change (e.g., a spike in cortisol or heart rate) influences both voice pitch and skin temperature, a multimodal system can use these redundant signals to increase the "Reliability Index" of its emotion recognition—a necessity for companion systems that operate in noisy, real-world environments.
Results & Key Focus Areas
The workshop identified several critical layers for the next generation of HCI:
- Timing: The importance of "when" a system responds in a multimodal scenario.
- Individuality: Incorporating personality traits into the emotion model to understand why two users might react differently to the same system error.
- Visualization: Using Self Assessment Manikins (SAM) to validate ground truth for Valence, Arousal, and Dominance.

Critical Analysis & Conclusion
Takeaway
The shift from "Emotion Recognition" (what is the user feeling?) to "Emotion Modelling for Companionship" (how does this feeling affect our long-term interaction?) is the workshop's greatest contribution. The emphasis on cognitive load (underload/overload) is particularly relevant for modern AI assistants that risk overwhelming users with information.
Limitations
While the 10-modal approach is robust, the computational overhead of processing such high-dimensional data in real-time remains a challenge for embedded companion devices. Furthermore, the dataset's focus on a specific age demographic (50+) may limit the generalizability of certain physiological markers to younger populations.
Future Outlook
As we move toward LLM-powered agents, the principles established in ERM4CT 2016 regarding interoperability and context-aware adaptation will be the foundation for moving AI from a "tool" to a "persona."
