Automated SNS Analysis: Merging SVM and Opinion Mining for Relational Insights
A Study on Automatic Analysis of Social Network Services Using Opinion Mining
This paper proposes an automated framework for analyzing user relationships on Social Network Services (SNS) by combining Support Vector Machines (SVM) for interaction classification and Opinion Mining for sentiment tendency prediction. The method specifically targets Twitter data, transforming complex user interactions into formalized feature sets to achieve efficient large-scale relational analysis.
TL;DR
As online communities shift from information-centered to relation-centered, the sheer volume of data makes manual analysis impossible. This paper introduces an automated pipeline that categorizes Twitter interactions using Support Vector Machines (SVM) and evaluates the emotional "tendency" of these interactions via Opinion Mining. The result is a scalable framework capable of identifying active user clusters and their relational sentiments.
The Shift from Data to Relations
The emergence of Social Network Services (SNS) has transformed the web into a person-to-person ecosystem. Unlike traditional forums, the value of an SNS lies in the strength and quality of interactions. However, researchers are often stuck using "classic approaches"—analyzing only fractions of data because they lack a formalized methodology to process millions of tweets, retweets, and mentions. This paper addresses the gap by asking: How can we automatically determine not just who talks to whom, but how they feel about that interaction?
Methodology: Formalizing the "Social" in SNA
The researchers break down the methodology into three critical phases: Feature Extraction, Structural Classification, and Semantic Analysis.
1. Feature Data Selection
The paper identifies five primary interaction types on Twitter: Tweet, Reply, Mention, DM, and Retweet. These are then mapped onto a feature matrix categorized by:
- Response: Measuring the "when" and "how" (reply-time, retweet-count).
- Relation: Measuring the "closeness" (direct vs. indirect followers).
- Interaction: The cumulative frequency of involvement.

2. SVM Classification
The core engine uses Support Vector Machines (SVM). Why SVM? Because it excels at handling non-linear problems by mapping data into higher-dimensional spaces. By using the "maximum margin classification" principle, the system draws a clear boundary between users who are "actively interacting" and those who are "isolated."

3. Opinion Mining Integration
Once users are grouped, the system applies Opinion Mining. This involves a pre-defined data dictionary that identifies positive and negative sentiments within those interactions. This layer adds a qualitative dimension to the quantitative structural data.
Experiments and Results
The authors tested their model on a dataset of 143 Twitter users. Using SVM Light, they purified the interaction data into a numerical format.
- Classification Success: The model effectively separated the user base into distinct "Active" and "Non-active" groups.
- Time Efficiency: By automating the process, the system proved capable of analyzing user tendencies in significantly less time than traditional SNA methods.

Critical Insight: Beyond the Graph
The true value of this work lies in its hybrid approach. While most Social Network Analysis (SNA) treats users as nodes and interactions as simple edges, this paper recognizes that an edge (an interaction) has a "color" (sentiment).
Limitations & Future Work
While the SVM approach is mathematically robust, the study's reliance on a static "data dictionary" for Opinion Mining may struggle with the evolving slang and sarcasm prevalent on Twitter. The authors acknowledge this, noting that future work will focus on building a more sophisticated automated tool and ensuring the "satisfiability and reliability" of the sentiment data.
Conclusion
By formalizing the "messy" data of human interaction into a machine-learnable feature set, Kwon and Park have provided a blueprint for real-time SNS monitoring. This moves us away from retrospective analysis toward predictive modeling of online community health.
