OLAPing Big Social Data: Bridging Multidimensional Analytics and Social Repositories
OLAPing Big Social Data: Multidimensional Big Data Analytics over Big Social Data Repositories
This paper introduces a holistic framework for "OLAPing Big Social Data," bridging the gap between traditional multidimensional analytics and semi-structured social media repositories. It proposes a modular reference architecture that leverages NoSQL technologies like MongoDB and big data ecosystems like Hadoop/Spark to enable interactive knowledge discovery over high-volume social datasets.
TL;DR
Social media data is a goldmine of business intelligence, yet it remains largely "locked" behind its semi-structured and massive-scale nature. This paper proposes a formal shift from simple procedural data processing to OLAPing Big Social Data—using multidimensional cubes to drive interactive, rich knowledge discovery. By integrating NoSQL storage with traditional OLAP operators, the author provides a roadmap for "actionable" social analytics at scale.
Background Positioning
In the spectrum of Big Data research, this work serves as an architectural synthesis and vision paper. It situates itself between the rigidity of classical Data Warehousing and the chaotic flexibility of the "Data Lake" approach, advocating for a middle ground where social data is semantically enriched into a "social cube."
The Core Problem: The Structure Gap
Why can't we just plug Facebook or Twitter data into a standard Business Intelligence (BI) tool?
- Schema Rigidness: Standard OLAP expects fixed hierarchies; social data is irregular and evolving.
- Unstructured Content: Most social value is buried in text (tweets, comments) which OLAP cannot natively "sum" or "average."
- Volume and Velocity: The "Big" in Big Social Data renders traditional relational databases obsolete for real-time multidimensional exploration.
Methodology: The Reference Architecture
The author outlines a modular, four-layer architecture designed to handle the "V's" of big social data while maintaining the analytical power of OLAP.
1. The Modular Stack
- Storage Layer: Moves away from SQL to MongoDB and Hadoop. The document-oriented nature of NoSQL is perfectly aligned with JSON-like social media objects.
- Analytics Layer: Employs Hive and Pentaho to translate big data into queryable structures.
- Knowledge Discovery Layer: The "Front-End" where Machine Learning (ML) meets human intuition through visual analytics.
2. Semantic Mapping (How it works)
To turn a tweet into a "measure," the framework suggests:
- Topic Modeling (LDA): To derive "Interest" dimensions.
- Sentiment Analysis: To create numerical "measures" (e.g., average sentiment per demographic).
Figure 1: The proposed modular reference architecture showing the flow from raw social sources to actionable knowledge.
Research Challenges: The Front Line
The author identifies several "hard problems" that define the current research frontier:
- OLAP-aware Graph Mining: Since social networks are essentially graphs, we need a "Graph Cube" where nodes and edges can be aggregated like sales figures in a spreadsheet.
- Privacy-Preserving OLAP: How do we perform multidimensional analysis on user behavior without compromising individual identities? The paper suggests probabilistic models as a potential solution.
- Web-scale Aggregation: Supporting
roll-upanddrill-downoperations across billions of social records in sub-second response times.
Experimental Case Studies
The paper references several state-of-the-art implementations that validate this architecture:
- Twitter Analysis [17]: Integrating text mining into warehousing to enrich tweet datasets semantically.
- MS-LDA Models [20]: Using
word2vecand social relationships to extract "hidden interests" as dimensions for OLAP. - Contextual Dimensions [21]: Automatically creating dimensions from social text using hierarchical clustering.
Critical Analysis & Conclusion
Takeaway: This paper effectively argues that social data analytics must evolve beyond "what happened" (descriptive) to "why and where" (multidimensional).
Limitations: While the architecture is theoretically sound and modular, the paper is a high-level survey/vision work. It lacks a singular, new benchmark result, instead relying on the synthesis of existing state-of-the-art methods. The "Integration with Modern Processing Platforms" (like Spark) is mentioned as a requirement but not detailed in terms of implementation specifics.
Future Outlook: We are moving toward a world of Visual Big Data Analytics. The future of this field lies in "dimensionally-parameterized" spaces where a business user can interactively browse a 10-billion-node social graph as easily as an Excel pivot table. Practitioners should focus on the "Document-to-Cube" mapping techniques discussed here.
