Thermal Face Analysis: Bridging the Gap Between Infrared and Visual Domains

A Thermal Infrared Face Database With Facial Landmarks and Emotion Labels

2018-12-20
Marcin Kopaczka, Raphael Kolk, Justus Schock, Felix Burkhard, Dorit Merhof
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a high-resolution thermal infrared face database (1024x768) featuring 2,935 images of 90 subjects with 68-point manual landmark annotations and emotion labels. The authors demonstrate that training data-driven models like Active Appearance Models (AAM) and Deep Alignment Networks (DAN) on this dataset significantly outperforms traditional rule-based thermal imaging methods.

TL;DR

Researchers from RWTH Aachen University have released a landmark high-resolution thermal infrared face database to solve the "data drought" in infrared imaging. By providing 2,935 images with meticulous 68-point annotations, they prove that modern data-driven methods like Deep Alignment Networks (DAN) and SVMs can finally bring thermal face detection and emotion recognition to a level of robustness once reserved for RGB cameras.

Background: Why Thermal Imaging is Hard

Thermal or Long-Wave Infrared (LWIR) imaging operates in the 7–14 µm wavelength. Unlike RGB cameras that capture reflected light, LWIR sensors capture emitted heat. This presents a unique challenge: texture and color cues—the bread and butter of traditional computer vision—are almost entirely absent.

While thermal imaging offers "superpowers" like working in total darkness and revealing physiological signals (heart rate, breathing), the algorithms have lagged behind. Most existing methods rely on simple thresholding (finding the hot face against a cold background), which fails miserably if the subject moves or changes pose.

The "Thermal Face Project" Methodology

The authors argue that the weakness isn't the thermal modality itself, but the lack of training data. To fix this, they created a database with:

  • High Resolution: 1024 x 768 pixels (standard databases are often 320x240).
  • Expert Annotation: 68 manual landmarks, following the Helen/LFPW standard.
  • Variety: 9 head pose sequences, 7 basic emotions, and specific Facial Action Units (AUs).

Architecture & Fitting

The study explores two primary paths for landmark detection:

  1. AAM with Pose Estimation: Using Random Forest regression to estimate head pose first, which then initializes an Active Appearance Model (AAM). This prevents the model from getting stuck in local minima during out-of-plane head rotations.
  2. Deep Alignment Network (DAN): A multi-stage Convolutional Neural Network (CNN) that refines landmark heatmaps across successive stages.

Model Architecture and Pose Estimation Flow Figure 1: The pipeline for head pose estimation and AAM initialization.


Experiments: Breaking the SOTA

The results prove that "borrowing" logic from the visual domain works if the data is there.

1. Landmark Detection Performance

The Deep Alignment Network (DAN) was the clear winner. Not only was it more precise, but it was also significantly faster due to its feedforward nature and GPU optimization.

  • DAN Speed: 33.3 FPS (Real-time)
  • AAM (HOG-WIC) Speed: ~0.14 FPS (Iterative and slow)

Fitting Results Comparison Figure 2: Lateral comparison between initialization methods. Note how the DAN (sixth column) closely matches the Ground Truth (last column) even in rotated poses.

2. Emotion Recognition

Can machines "feel" the heat? The researchers tested several classifiers (SVM, kNN, Random Forest) against various feature descriptors (HOG, LBP, DSIFT).

  • The Winner: Linear SVM + DSIFT.
  • Key Finding: In an 8-emotion classification task, the algorithm achieved 46.7% accuracy, actually beating human observers who averaged 42.0%. Humans tend to over-classify ambiguous thermal faces as "neutral," whereas the SVM captures subtle thermal variations.

Depth Insight: Why It Matters

This research shifts the paradigm for thermal imaging. Previously, scientists spent years developing custom "thermal-only" rules. This paper demonstrates that transferring the Inductive Bias from visual domain architectures (like DAN or HOG-SVM) is a more viable path, provided we invest in high-fidelity annotations.

Limitations & Future Work

  • Glasses: The current database excludes subjects wearing glasses (which are opaque in thermal). Future iterations must handle the "black hole" effect glasses create.
  • Background: The tests were done against neutral backdrops. Testing "in-the-wild" with complex thermal environments (radiators, sun-heated walls) is the next frontier.

Summary

By marrying high-resolution thermal sensors with deep learning, the "Thermal Face Project" provides a blueprint for bringing infrared computer vision into the modern era. The database is available for academic use, inviting the community to further refine how we process the "invisible" spectrum.

Explore the data at: https://github.com/marcinkopaczka/thermalfaceproject

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Deep Alignment Networks (DAN) or similar architectures for cross-modal face recognition between RGB and thermal infrared domains.
  • What are the current SOTA methods for "in-the-wild" thermal facial landmark detection that handle unconstrained backgrounds and varying sensor focal depths?
  • Identify studies that apply thermal facial landmark tracking to clinical monitoring tasks, such as non-contact sleep apnea detection or respiratory rate estimation.
Contents
Thermal Face Analysis: Bridging the Gap Between Infrared and Visual Domains
1. TL;DR
2. Background: Why Thermal Imaging is Hard
3. The "Thermal Face Project" Methodology
3.1. Architecture & Fitting
4. Experiments: Breaking the SOTA
4.1. 1. Landmark Detection Performance
4.2. 2. Emotion Recognition
5. Depth Insight: Why It Matters
5.1. Limitations & Future Work
6. Summary