Beyond Deep Learning: Robust Emotion Recognition via Optimized Logistic Regression
Emotion Recognition Through Facial Expressions Using Supervised Learning with Logistic Regression
2021-01-01
Summary
Problem
Method
Results
Takeaways
Abstract
The paper presents a facial emotion recognition system based on Logistic Regression and Supervised Learning to classify seven universal emotions plus a neutral state. By leveraging the dlib library for 68 landmark extraction and a "one-vs-all" classification strategy, the model achieves a competitive 85% accuracy on the RaFD dataset, performing on par with commercial APIs like Microsoft Face and Face++.
## TL;DR
While Deep Learning dominates Computer Vision, this study proves that a carefully optimized **Logistic Regression** model can achieve **85% accuracy** in emotion recognition. By extracting 68 facial landmarks and applying rigorous mathematical normalization, this approach matches the performance of heavyweight industry solutions like **Microsoft Face** and **Face++**.
## Background & Motivation: The Power of Facial Expression
Research suggests that up to **55% of message meaning** is conveyed through facial expressions rather than words. In the realm of **Affective Computing**, bridging the gap between human emotion and machine understanding is vital. The authors set out to build a system capable of recognizing the seven universal emotions defined by Paul Ekman’s theory: Anger, Contempt, Disgust, Fear, Happiness, Sadness, and Surprise.
## Methodology: The "One-vs-All" Strategy
The authors utilized the **Radboud Faces Database (RaFD)**, known for high-resolution images and controlled lighting. The technical pipeline follows three main stages:
1. **Landmark Extraction**: Using the **dlib** library, the system identifies 68 specific "landmarks" on the face (eyes, brows, nose, mouth, and contour), resulting in a 136-element characteristic vector ($x, y$ coordinates).
2. **Logistic Regression**: Since emotion recognition is a multi-class problem, a **"one-vs-all"** method was used. Eight separate classifiers were trained, where each calculates the probability of an image belonging to a specific emotion class.
3. **Advanced Optimization**: The $fminunc$ (unconstrained function minimization) algorithm was employed to minimize the cost function $J( heta)$.

*Fig 1: The operational sequence from landmark extraction to final emotion prediction.*
## The Secret Sauce: Data Normalization
One of the most technical insights of the paper is the impact of **Feature Scaling** and **Mean Normalization**. Without these, the cost function gradients can become "distorted," making convergence extremely slow.
Using the formula:
$$x_i := \frac{x_i - \mu_i}{s_i}$$
(where $\mu$ is the mean and $s$ is the standard deviation), the authors forced all characteristics into a tight range between -5 and 5. As shown in the training logs, **normalized data reached a local minimum significantly faster**, whereas unnormalized data often hit the 1200-iteration limit without converging.
## Experimental Battle: S.P. vs. Tech Giants
The authors didn't just test internally; they pitted their "System Proposed" (S.P.) against **Microsoft Face** and **Face++**.
### Key Performance Metrics:
* **Accuracy**: 85% (Global).
* **Confusion Matrix Insights**: The model was perfect at identifying **Disgust, Happiness, and Surprise** (Recall = 1.0). Its main weakness was "Contempt," which it occasionally confused with "Sadness" or "Neutral."
* **Comparison**: On the CK+ dataset, the S.P. made only **2 errors**, identical to Face++ and slightly trailing Microsoft. Interestingly, Microsoft Face **consistently failed** at detecting "Contempt," a gap where the S.P. showed higher sensitivity.

*Fig 2: Confusion Matrix showing high precision in positive emotions but challenges in subtle "Contempt" detection.*
## Critical Analysis & Future Outlook
The study demonstrates that **Logistic Regression** is not obsolete. It offers a transparent, computationally efficient alternative to black-box neural networks.
**Limitations**: The model performs slightly worse on low-resolution images (like those from web snippets) compared to high-res studio shots. This is likely due to the landmark detector's sensitivity to pixel density.
**Future Work**: The authors plan to integrate this facial recognition model into **multimodal systems**, combining it with heart rate sensors and EEG-based brain-computer interfaces to create a more holistic "emotional profile" of the user. This paves the way for advanced applications in healthcare, education, and human-robot interaction.
