Skl-MHI: Efficient Human Action Recognition through Skeleton Temporal Compression
Skeleton motion history based human action recognition using deep learning
This paper introduces a Human Action Recognition (HAR) system that combines Skeleton Motion History Images (Skl MHI) with a 2D Deep Convolutional Neural Network (2D-DCNN). By encoding temporal dynamics into a single 2D skeleton-based representation, the method achieves a high recognition accuracy of 97.82% across 10 distinct action categories using Kinect V2 sensor data.
TL;DR
The paper presents a streamlined approach to Human Action Recognition (HAR) by transforming 3D skeleton sequences into 2D Skeleton Motion History Images (Skl MHI). By training a lightweight 2D-DCNN on these compressed representations, the researchers achieved a 97.82% accuracy while significantly reducing the computational burden compared to traditional raw-video deep learning.
Problem & Motivation: The High Cost of Action Recognition
Recognizing human actions—like walking, bending, or waving—is crucial for elderly monitoring and human-robot interaction. However, the field has long struggled with two primary hurdles:
- Handcrafted Frailty: Older methods depend on manual feature extraction tailored to specific environments, making them brittle when background or lighting changes.
- Computational Complexity: Directly processing 3D video volumes (RGB-T) requires massive memory and GPU power.
The authors' insight is to leverage the Microsoft Kinect V2 to extract skeleton joint points first. This removes environmental noise (background clutter) and reduces the data dimensionality from millions of pixels to just 25 joint coordinates, focusing purely on human kinematics.
Methodology: From Skeleton Sequences to Skl MHI
The core innovation lies in how temporal information is packed into a static image.
1. The Skl MHI Generation
Instead of feeding a sequence of frames, the system takes 9 consecutive skeleton frames. It performs a binary OR-operation across these frames to create a "motion trail" or a history of the movement. This result is normalized into a 62×62 image, creating a signature of the action's trajectory.

2. Lightweight 2D-DCNN Architecture
Because the input is a compact 62x62 image, the neural network doesn't need to be deep. The proposed architecture uses:
- 3 Convolutional Layers: Utilizing Gabor and Gaussian filters for feature extraction.
- Max Pooling: For spatial invariance.
- Soft-max Output: To classify the 10 specific actions (e.g., A1: Bending, A6: Waving).
Experiments & Results: Precision Meets Efficiency
The model was tested on a self-collected dataset featuring 10 actions performed by 6 different individuals.
SOTA Comparison and Accuracy
The system showed remarkable convergence speed. It hit 94.52% accuracy in just 10 epochs. When pushed to 1000 epochs, the accuracy stabilized at 97.82%.

As shown in the confusion matrix, actions like "Standing" (A4), "Pointing" (A9), and "Lying" (A10) achieved 100% accuracy, demonstrating that the Skl MHI captures unique geometric signatures for these postures perfectly. Some confusion was noted between "Bending" (A1) and "Taking Medicine" (A5) due to similar torso inclinations.

Critical Analysis & Conclusion
Takeaway
The beauty of this research is its simplicity. By converting a temporal problem into a spatial one through Skl MHI, the authors circumvent the need for expensive Recurrent Neural Networks (RNNs) or 3D-CNNs. This makes the system ideal for deployment on edge devices with limited processing power.
Limitations & Future Work
The current approach relies on a fixed viewpoint. If the person turns 90 degrees, the 2D "silhouette" of the skeleton changes significantly. The authors acknowledge this and plan to investigate view-invariant environments and more diverse action sets in future iterations. To truly move HAR forward, integrating this Skl MHI approach with Graph Convolutional Networks (GCNs) could further exploit the topological structure of the human skeleton.
