Beyond the Questionnaire: Merging Thai Social Text Mining with Lifestyle Segmentation

Incorporating Social Network Thai Text Mining with Lifestyle Segmentation Analysis

2017-07-01
Nitipan Ratanasawadwat, Rachsuda Jiamthapthaksin
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a hybrid marketing framework that combines lifestyle segmentation through online surveys (CASI) with Thai text mining of Facebook posts. By integrating k-means clustering and Latent Dirichlet Allocation (LDA), it achieves a more nuanced profile of young adult consumers in Thailand, categorizing them into four distinct psychographic segments.

TL;DR

Marketing segmentation is evolving from static surveys to dynamic behavioral analysis. This paper proposes a dual-track framework that combines traditional Computer-Assisted Self-Interview (CASI) questionnaires with Thai text mining from Facebook. By mapping lifestyle factors (like Materialism and Self-Confidence) against social media topics (extracted via LDA), the authors identify unique consumer archetypes such as "Social-Materialists" and "Self-Esteem" seekers.

Background Positioning

In the landscape of market research, this study serves as a bridge between traditional psychometrics and modern Big Data analytics. It targets the "always-on" generation of Thai young adults, using the Facebook ecosystem as a laboratory for genuine consumer insight.

The Problem: The Bias of Self-Reporting

Traditional surveys often face the Social Desirability Bias, where respondents answer in a way they think they should behave. Additionally, manual data entry is slow and error-prone. The authors argue that while people might "filter" their survey answers, their Facebook posts act as a more liberal expression of their anxieties, interests, and daily routines, governed by Self-Regulation Theory (SRT).

Methodology: The Fusion of Two Worlds

The methodology is a rigorous pipeline that transforms unstructured Thai text and structured survey data into a unified clustering space.

1. The Data Pipeline

The framework utilizes a custom Facebook application to crawl user posts via the Graph API, while simultaneously administering a 24-item AIO (Activities, Interests, Opinions) lifestyle survey.

Data Pre-processing Framework

2. Thai Text Pre-processing

Thai language presents unique challenges, specifically the lack of word boundaries. The study employs:

  • LongLexTo: A longest-matching tokenization technique.
  • TF-IDF: To filter out common stop words and identify high-value keywords.
  • Latent Dirichlet Allocation (LDA): To group keywords into 10 distinct "Topics of Interest" (e.g., Sports, Shopping, Love, Religion).

3. Psychographic Factor Analysis

The 24 AIO items are compressed into 6 core factors using Principal Component Analysis:

  1. Materialism
  2. Self-Confidence
  3. Planner
  4. Disciplinary
  5. Self-concerning
  6. Socialism

Experiments and Results: The Four Consumer Archetypes

By applying k-means clustering to the combined dataset of lifestyle factors and LDA topics, four distinct personas emerged:

ClusterDefining TraitsFacebook Content Focus
Self-esteemHigh Confidence & DisciplineFrequently posts about "Life" and "Love"
Self-concernedHigh Planning, Low MaterialismModerate focus on "Life" routines
Self-isolatedLow Confidence & Social interestLeast likely to post about "Love"
Social-materialismHigh Materialism & Social activityInterestingly, posts more about "Religion" than personal life

Cluster Analysis Results

Key Insight: The "Social-Materialism" Paradox

One of the most intriguing findings is that the "Social-Materialism" group avoids posting about their personal lives, opting instead for "Religion" topics. This suggests that for these individuals, social media is a platform for public image management rather than personal vulnerability—a classic example of the Self-Regulation Theory in action.

Critical Analysis & Conclusion

Takeaway

The integration of social media text mining doesn't just add more data; it adds context. Knowing that a consumer is "Materialistic" is one thing; knowing they express this through "Religion/Temple" posts on Facebook allows for a radically different marketing approach (e.g., CSR-focused branding vs. direct luxury ads).

Limitations

  • API Volatility: The study relies heavily on the Facebook Graph API, which has become significantly more restrictive regarding user data access in recent years.
  • Sample Homogeneity: The sample (students from Assumption University) may not represent the broader Thai population.

Future Outlook

Future research should look toward Sentiment Analysis (analyzing how they feel, not just what they talk about) and expanding the framework to platforms like TikTok or Twitter, where Thai linguistic nuances (slang and short-form text) are even more prevalent.

Find Similar Papers

Try Our Examples

  • Search for recent papers that integrate Latent Dirichlet Allocation (LDA) with k-means clustering for psychographic market segmentation.
  • Which study first introduced the use of Facebook Graph API for behavioral research, and how has API privacy restriction changed this methodology since 2017?
  • Explore how Thai natural language processing (NLP) techniques like LongLexTo have been adapted for sentiment analysis in modern social commerce.
Contents
Beyond the Questionnaire: Merging Thai Social Text Mining with Lifestyle Segmentation
1. TL;DR
2. Background Positioning
3. The Problem: The Bias of Self-Reporting
4. Methodology: The Fusion of Two Worlds
4.1. 1. The Data Pipeline
4.2. 2. Thai Text Pre-processing
4.3. 3. Psychographic Factor Analysis
5. Experiments and Results: The Four Consumer Archetypes
5.1. Key Insight: The "Social-Materialism" Paradox
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook