EmoGANs: Synthesizing the Complexity of Human Emotion through Deep Adversarial Learning

Generation of Compound Emotions Expressions with Emotion Generative Adversarial Networks (EmoGANs)

2020-09-23
Win Shwe Sin Khine, Prarinya Siritanawan, Kazunori Kotani
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces EmoGANs, a generative framework based on DCGANs designed to synthesize compound facial expressions (e.g., "happily surprised") by blending features from basic emotion categories. It achieves effective emotion manipulation using a multi-input generator architecture and semi-supervised training on the CK+ and JAFFE datasets.

TL;DR

Human feelings are rarely "pure." While AI has mastered the six basic emotions (Happy, Sad, etc.), it struggles with "compound emotions" like being "happily disgusted." EmoGANs (Emotion Generative Adversarial Networks) bridges this gap by blending basic facial features into complex expressions using a novel multi-input DCGAN architecture and similarity-based image morphing.

Beyond the Six Basics: The Motivation

For decades, the field of facial expression recognition has been anchored by Ekman’s six universal basic emotions. However, human experience is nuanced. A graduation ceremony, for instance, produces a "bittersweet" expression—a hybrid of happiness and sadness.

The current bottleneck is data: humans are poor at "faking" compound emotions on command for datasets. Authors Win Shwe Sin Khine et al. realized that instead of asking humans to act, we could teach machines to synthesize these expressions by understanding the underlying manifold of facial muscle movements.

Methodology: The EmoGANs Architecture

The core innovation lies in the transition from a traditional GAN to a multi-input structure.

1. Feature-Concatenated Generation

Standard GANs map a random noise vector () to an image. EmoGANs, however, feeds the generator two base images ( and ) representing basic emotions.

  • Block 1: Extracts prominent facial features via strided convolutions.
  • Concatenation: Merges these features with latent noise to form a "compound feature vector."
  • Upsampling: Uses transposed convolutions to reconstruct a 64x64 image reflecting the mixed state.

EmoGANs Generator Architecture

2. The Cosine Similarity Trick

To train the model, the authors needed "Pseudo-Ground Truth." They used image morphing to blend two basic emotions. To ensure the quality of these morphed images, they employed Cosine Similarity to select pairs that were structurally compatible but emotionally distinct, preventing the "ghosting" effects common in random image blending.

Experiments and Visual Evidence

The authors tested EmoGANs on the CK+ and JAFFE datasets. A critical finding was the superiority of semi-supervised learning.

Unsupervised vs. Semi-Supervised

Initial unsupervised attempts resulted in mode collapse or blurry outputs because the discriminator struggled to define what a "correct" compound emotion looked like. By introducing labels (Semi-Supervised), the discriminator acted as a classifier, forcing the generator to produce features recognizable as both constituent emotions.

Comparison of Training Results Fig: Training curves showing the convergence of loss and accuracy, stabilizing the internal representation of complex facial features.

As seen in the results generated from the JAFFE dataset, the model successfully captures the subtle overlapping of muscle movements (e.g., eye narrowing from "anger" combined with mouth curvature from "surprise").

Generated Compound Emotions Fig: Synthesized compound expressions. Note the retention of identity while altering the emotional blend.

Critical Insight: Why This Matters

The technical takeaway here is the validation of the Linear Combination Model for emotions. By treating emotions as vectors in a latent space that can be summed and weighted, EmoGANs proves that we can "compute" complex human states.

Limitations: The current resolution is 64x64, which limits the fine-grained representation of micro-expressions. Additionally, while the model balances two emotions, real human states might involve three or more simultaneous signals.

Conclusion

EmoGANs represents a significant step toward more empathetic AI. By moving beyond categorical "either/or" emotion labels to a fluid, generative "both/and" approach, this research opens the door for virtual avatars and interaction systems that truly mirror the complexity of human cognition.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use StyleGAN or Diffusion Models to generate fine-grained or compound facial expressions beyond the basic six categories.
  • Which paper first established the mathematical definition of facial compound emotions using Action Units (AUs), and how does the linear combination model in this paper differ?
  • Explore how generative models for facial expressions are being integrated into real-time Human-Computer Interaction (HCI) systems for virtual learning or empathetic AI.
Contents
EmoGANs: Synthesizing the Complexity of Human Emotion through Deep Adversarial Learning
1. TL;DR
2. Beyond the Six Basics: The Motivation
3. Methodology: The EmoGANs Architecture
3.1. 1. Feature-Concatenated Generation
3.2. 2. The Cosine Similarity Trick
4. Experiments and Visual Evidence
4.1. Unsupervised vs. Semi-Supervised
5. Critical Insight: Why This Matters
6. Conclusion