iOLAP: Bridging the Gap Between Database OLAP and the Unstructured Web

iOLAP: A Framework for Analyzing the Internet, Social Networks, and Other Networked Data

2009-03-19
Yun Chi, Shenghuo Zhu, Koji Hino, Yihong Gong, Yi Zhang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces iOLAP, a systematic framework for analyzing high-dimensional networked data (Internet, social networks, citations) using a Polyadic Factor Model. By representing data as a multi-dimensional "networked data cube" and employing Non-negative Tensor Factorization (NTF), the framework enables OLAP-style operations like roll-up and drill-down to uncover latent structures across People, Relation, Content, and Time dimensions.

TL;DR

The paper introduces iOLAP, a framework that applies the multidimensional analysis rigor of traditional Business Intelligence (OLAP) to the messy, linked world of the Internet. By using a Polyadic Factor Model (based on Tensor Factorization), it treats People, Relations, Content, and Time as inter-dependent dimensions. It achieves superior personalization in recommendations and deep exploratory insights into social communities while maintaining linear scalability.

Background: Beyond the Search Box

For decades, we have interacted with the web primarily through search engines. However, search engines are designed for retrieval—finding a needle in a haystack. They follow a power-law distribution, favoring dominant sites and ignoring the "long tail" where grassroots sentiments and niche wisdom reside.

The authors argue that just as databases evolved from simple SQL queries to OLAP (Online Analytical Processing) for business intelligence, the Internet needs an Internet OLAP to visualize distributions, summarize trends, and analyze high-order relationships across heterogeneous data.

The Core Innovation: Polyadic Factorization

The biggest challenge in "Internet OLAP" is that networked data is multi-modal. A single event involves a person (People), a link (Relation), a message (Content), and a timestamp (Time).

Why pairwise isn't enough

Most prior works used a "two-step" approach:

  1. Analyze People vs. Content.
  2. Analyze People vs. Relations.
  3. Try to glue the results together.

This misses the joint effect. For instance, a specific topic might only be discussed by a specific group of friends during a specific week. iOLAP solves this by using Polyadic Factorization, which models all dimensions simultaneously.

Mathematical Intuition

The framework utilizes a non-negative version of the Tucker Decomposition. It decomposes a large, sparse data tensor into a smaller "core tensor" and factor matrices () for each dimension.

Model Architecture

The core tensor is the "brain" of the model—it captures how the latent groups in one dimension correlate with groups in others.

Scalability through "Lazy Computation"

Tensor operations are notoriously expensive (). To make iOLAP practical for millions of records, the authors developed an optimization that:

  1. Exploits Sparseness: Only computes values for non-zero entries.
  2. Ordered Computation: Reuses matrix products to minimize redundant floating-point operations. The result is an algorithm that is linear with respect to the number of records, making it feasible for massive datasets like CiteSeer.

Experimental Insights: Blogs and Citations

The authors applied iOLAP to two distinct domains:

1. The Blogosphere

By "rolling up" the core tensor, iOLAP identified that different blog groups might share a topic (e.g., Politics), but with different nuances (e.g., one group focusing specifically on terrorism).

Community Visualization Figure: Time dimension factors showing temporal patterns of blog activity.

2. Personalized Recommendation (CiteSeer)

Standard recommendations suggest "popular" papers. iOLAP allows for triple-conditional recommendation: .

The results were striking: iOLAP outperformed popularity-based baselines significantly.

MethodTop-1 NDCGTop-10 NDCG
Popularity Baseline0.0250.076
iOLAP (Factor_KL_10)0.0380.120

Critical Insight & Conclusion

The true power of iOLAP lies in its interpretability. Unlike black-box deep learning models, the factor matrices and core tensors provide a readable map of how communities, topics, and time intersect.

Limitations: The current framework captures the intensity of communities but doesn't easily track the structural evolution (e.g., when one community splits into two).

Future Work: The authors envision expanding this to automated summarization and more complex OLAP operations to provide a full "Business Intelligence" suite for the decentralized web.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Non-negative Tensor Factorization (NTF) to dynamic or streaming networked data to capture community evolution.
  • What are the theoretical links between Probabilistic Latent Semantic Analysis (PLSA) and Tucker-style tensor decompositions in multi-way clustering?
  • Explore how the iOLAP framework's "People-Relation-Content-Time" dimensions are adapted in modern Graph Neural Networks (GNNs) for heterogeneous information networks.
Contents
iOLAP: Bridging the Gap Between Database OLAP and the Unstructured Web
1. TL;DR
2. Background: Beyond the Search Box
3. The Core Innovation: Polyadic Factorization
3.1. Why pairwise isn't enough
3.2. Mathematical Intuition
4. Scalability through "Lazy Computation"
5. Experimental Insights: Blogs and Citations
5.1. 1. The Blogosphere
5.2. 2. Personalized Recommendation (CiteSeer)
6. Critical Insight & Conclusion