Decoding the Instagram Fashionista: A Data Mining Approach to Customer Behavior
Analysis of the behavior of customers in the social networks using data mining techniques
This paper presents a data mining framework using the CRISP-DM methodology to analyze customer behavior on Instagram for a fashion company. By applying K-means clustering and FP-Growth association rules, the study successfully identifies high-engagement product categories like "Dressing Gowns" and "Long Cocktail Dresses" to inform B2C marketing strategies.
TL;DR
In the fast-paced world of fashion, understanding consumer "likes" is more than a vanity metric—it's a goldmine for strategy. This paper leverages the CRISP-DM methodology and K-means clustering to dissect Instagram data from a leading fashion brand, identifying which specific clothing styles (like cocktail dresses) drive the highest engagement and how association rules can predict future marketing success.
Background: Beyond the Feed
As companies shift toward Social CRM, the goal has moved from simple broadcasting to personalized engagement. The authors pose a powerful hypothesis: The predictions of future customer behavior are rooted in the past behaviors of others. By analyzing interactions on Instagram, businesses can move away from guesswork and toward a scientific understanding of user preferences.
Problem & Motivation: The Data-Information Gap
The digital era presents a paradox: companies are drowning in data but starving for insights. Many fashion brands post content based on intuition rather than empirical evidence. The specific challenge addressed here is characterizing which attributes (media type, tags, filters) actually trigger a customer to move from passive scrolling to active engagement (liking/commenting).
Methodology: The CRISP-DM Framework
The authors adopted the CRISP-DM (Cross Industry Standard Process for Data Mining) workflow, ensuring a structured transition from business understanding to deployment.
1. Data Acquisition & Preparation
Using the Instagram API and Python, the researchers gathered 1,435 records. Key attributes included creation date, media type, filter type, likes, comments, and tags.
2. The Modeling Core
The study implemented two primary descriptive techniques using RapidMiner:
- Clustering (K-means): To segment the data into groups with similar interaction profiles.
- Association Rules (FP-Growth): To find frequent relations between different attributes (e.g., how "media type" relates to "likes").
Above: Table 1 reveals the different clusters based on engagement metrics (Likes/Comments).
Experiments & Results: What Makes a Post Viral?
The analysis yielded specific, actionable segments. By evaluating the Sum of Squared Errors (SSE), the authors determined that K=5 provided the optimal number of clusters.
- Cluster 0 (High Impact): Focused on "Dressing Gowns," showing high engagement (200-250 likes).
- Cluster 4 (The Powerhouse): Focused on "Long Cocktail Dresses." This was identified as the company’s core strength, with 299 images generating peak interaction levels.
- Temporal Insights: Association rules (via FP-Growth) indicated that posts from 2012-2013 had specific characteristics that led to historical engagement highs, suggesting a need to revisit those content styles.
Above: The FP-Growth results showing the support and size of frequent itemsets.
Deep Insights & Conclusion
Takeaways for the Industry
The study proves that unsupervised learning is an efficient alternative to traditional market research. By identifying that "Long Cocktail Dresses" are the primary engagement driver, the brand can optimize its inventory and marketing spend toward these high-performing assets.
Limitations & Future Work
While effective, the study is limited to numerical and categorical metadata. A significant future expansion would be to incorporate Computer Vision (CNNs) to analyze the visual features of the clothing itself—such as color, texture, and pattern—to see how they influence the K-means clustering. Additionally, expanding the scope to Facebook and TikTok would allow for a cross-platform understanding of customer "social personas."
Final Summary
This paper serves as a practical roadmap for SMEs in the fashion industry to transition into data-driven powerhouses, proving that even simple data mining techniques can yield significant competitive advantages in the B2C landscape.
