GRMAC: Winning the "AI Meets Beauty" Challenge with Flexible Attention
Beauty Product Retrieval Based on Regional Maximum Activation of Convolutions with Generalized Aention
This paper introduces Generalized-attention Regional Maximum Activation of Convolutions (GRMAC), a novel image descriptor tailored for beauty and personal care product retrieval. By integrating a flexible attention mechanism with multi-model feature fusion, the method achieved 1st place in the ACM Multimedia 2019 "AI Meets Beauty" Grand Challenge with a MAP@7 score of 0.4086.
TL;DR
Determining the "right" product in a database of half a million beauty items is notoriously difficult due to cluttered backgrounds and subtle packaging differences. The USTC NELSLIP team secured 1st place at ACM MM 2019 by introducing GRMAC, a descriptor that uses a tunable hyperparameter to "squeeze" out background noise, paired with a strategic fusion of global and regional deep features.
Background: Why Beauty Product Retrieval is Hard
In traditional image retrieval (like landmarks), the geometry is often rigid. In beauty products, however, a bottle of lotion might look entirely different depending on the camera angle, or it might be surrounded by dozens of other items on a shelf.
Previous SOTA methods like RMAC (Regional Maximum Activation of Convolutions) extract features from various local regions and sum them up. The flaw? It treats the chaotic background and the actual product with the same level of importance.
Methodology: The Power of the Generalized Attention Mask
The core innovation is the Generalized-attention Regional Maximum Activation of Convolutions (GRMAC).
1. Beyond Mean Thresholding
Earlier attention-based descriptors used the "mean value" of activations to decide which regions to keep. However, mean values are sensitive to outliers. GRMAC introduces a hyperparameter to calculate a more robust threshold : By tuning , researchers can control how "strict" the filter is. As increases, the mask becomes more selective, focusing only on the most salient parts of the product.
2. The Framework
The team utilized a dual-backbone approach:
- DenseNet201: Known for its feature reuse capabilities.
- SE-ResNet152: Leveraging Squeeze-and-Excitation blocks to weigh channel-wise importance.
Figure 1: The GRMAC pipeline—from raw image to fused feature vector.
Experiments & Results
The authors tested their method on the Perfect-500K dataset.
The Impact of Hyperparameter
A critical discovery was that the optimal value of varies by model. For SE-ResNet architectures, a specific effectively "cleans up" the mask, whereas a (equivalent to the older RAMAC method) often includes too much background noise.
Figure 2: Visual proof—as p increases (left to right), the attention mask focuses more tightly on the product.
Feature Fusion: The Winning Formula
The team found that Feature Fusion is the secret sauce. Specifically, combining a Global descriptor (MAC) with a Regional descriptor (GRMAC) yielded the best results. This suggests that while regional features capture fine-grained details, global features provide necessary context that prevents the model from "over-focusing" on small parts of the packaging.
| Descriptor Pair | Result (MAP) |
|---|---|
| DenseNet201 (GRMAC) + SE-ResNet152 (MAC) | 0.3877 (Validation) |
| Final Competition Score (Test) | 0.4086 |
Critical Analysis & Conclusion
Takeaway
GRMAC proves that attention in image retrieval doesn't have to be a complicated "black box" transformer layer. Sometimes, a mathematically sound, tunable thresholding mechanism applied to convolutional feature maps is more effective and computationally efficient for large-scale indexing.
Limitations & Future Work
While the method won the competition, the hyperparameter currently requires manual tuning (Ablation). A logical next step would be an Adaptive GRMAC where is learned per image or per category using a small policy network, allowing the model to decide how much background suppression is needed based on the image's complexity.
Author Insight: This work highlights a shift in CBIR (Content-Based Image Retrieval)—moving away from simple global pooling toward "Smarter Aggregation" that understands object-background relationships.
