Breaking the Silos: Semantics-Based Business Categorization for Multi-Platform Social Data
A Semantics-Based Approach for Business Categorization on Social Networking Sites
This paper introduces a semantics-based approach for cross-platform business categorization on Social Networking Sites (SNSs). The method, implemented in the CoDiT (Company Discovery Tool), uses semantic and morphological expansion of textual metadata to map heterogeneous SNS profiles (like Facebook and LinkedIn) into a standardized category system.
TL;DR
The proliferation of business pages across Facebook, LinkedIn, and other Social Networking Sites (SNSs) has created "data islands" with incompatible categorization schemes. This paper proposes a semantic-based methodology to unify these disparate categories by analyzing the textual content of business profiles. By leveraging semantic expansion and cosine similarity, the proposed CoDiT (Company Discovery Tool) achieves a significant leap from generic tags to precise, actionable business classifications.
The Motivation: Why "Company" is Not a Category
In the current digital ecosystem, a business might be an "Industry: Technology" entity on LinkedIn but simply a "Local Business" on Facebook. For developers building integrated search tools, these platform dependencies are a nightmare for two main reasons:
- Vagueness: Generic tags like "Organization" provide zero value for specific filtering.
- Schema Mismatch: Manual mapping between 100+ categories per platform is unscalable and fails whenever a platform updates its UI or API.
The authors' insight is simple yet powerful: the answer is in the text. Instead of relying on the platform's selected tag, we should analyze the "About" and "Products" sections to determine what a business actually does.
Methodology: The Three-Step Semantic Bridge
The core of the methodology is a pipeline designed to transform a simple category label into a robust search query that can be matched against a business description.
1. Preprocessing and Tokenization
Raw category tags (e.g., "Computer Graphics") are normalized and stripped of stop words to ensure the core semantic tokens are isolated.
2. Semantic and Morphological Expansion
This is the "secret sauce." To bridge the vocabulary gap between a category name and a marketing description:
- Semantic Expansion: Using the Leipzig Corpora Collection (LCC) API, the system finds synonyms and related terms (e.g., "Software" linked to "Application").
- Morphological Expansion: Using the
phpMorphylibrary, the system generates all lexical forms (singular/plural, tenses) to ensure a match regardless of sentence structure.

3. Cosine Similarity Measurement
The system treats the expanded category as a "bag of words" and the business profile as another. It then calculates the Cosine Similarity to determine the degree of overlap in a high-dimensional vector space.
Experimental Results: From One Tag to Thirteen
The system was tested on a dataset of Facebook business descriptions. The results were impressive, maintaining high performance during manual verification:
- Recall: 0.85
- F-measure: 0.79
The real-world value is best illustrated by the IBM Case Study. While Facebook's API returned only one category ("Company"), the CoDiT tool successfully extracted 13 relevant categories including Consulting, Research, Engineering, and Computer Hardware.

Critical Analysis & Conclusion
Takeaway
This research provides a robust framework for data integration in the Web 2.0 era. By moving from explicit labels to implicit semantic analysis, it becomes platform-agnostic, meaning it can theoretically integrate any new SNS (like TikTok for Business or Xing) with minimal adjustment.
Limitations & Future Work
The approach's accuracy is heavily dependent on the volume of text provided in the profile. A business with a blank "About" section remains unclassifiable. Future research could investigate using Computer Vision to analyze profile pictures or "Header" images to supplement missing textual data, or employ LLM-based Zero-shot classification for even higher precision in ambiguous cases.
Ultimately, this paper serves as a blueprint for building intelligent, unified discovery tools in an increasingly fragmented social world.
