MOR Model: Bridging the Gap Between Linked Open Data and Opinion Mining

A New Approach to Ontology-Based Semantic Modelling for Opinion Mining

2016-04-01
Rowida Alfrjani, Taha Osman, Georgina Cosma
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel Semantic Knowledge-Based (SKB) methodology for opinion mining, specifically targeting feature-level sentiment analysis in movie reviews. It presents the MOR (Movie, Opinion, and Review) model, which integrates domain-specific ontologies with automated enrichment from Linked Open Data (LOD) sources like DBpedia.

TL;DR

The paper presents a new methodology for semantic modelling in opinion mining that transcends simple keyword matching. By creating the MOR (Movie, Opinion, and Review) framework, the authors leverage Linked Open Data (LOD) to automatically populate a formal ontology with real-world facts. This structured knowledge is then integrated into a Natural Language Processing (NLP) pipeline to achieve high-precision feature-level sentiment analysis.

The "Knowledge Gap" in Sentiment Analysis

Most sentiment analysis tools operate on a "shallow" level—counting positive and negative words or using statistical patterns. These methods often fail at feature-level mining, where the goal is to understand what specifically a person liked (e.g., the "cinematography" vs. the "acting").

Prior works suffered from two main limitations:

  1. The Frequency Trap: Association Rule Mining (ARM) only finds features that appear often, ignoring niche but vital feedback.
  2. The Manual Bottleneck: Machine Learning (ML) approaches require a massive amount of human-labeled data, which is time-consuming and domain-specific.

The authors' insight was to use the Semantic Web as a shortcut. Instead of teaching a model what an "actor" is, why not link to a global database that already knows every actor's name?

Methodology: The MOR Framework

The core of the research is the MOR Model, which defines how data from three different domains interact:

  1. Movie Domain: Concepts like Star, Writer, Director, and technical features like Special Effects or Pacing.
  2. Opinion Domain: Utilizing the Marl Ontology to standardize how sentiment orientation (polarity) is recorded.
  3. Review Domain: Handling the metadata of the source text (Review ID, Reviewer Name).

Architecture and Enrichment

One of the most innovative parts of this workflow is the Automated Enrichment Process. Instead of manually entering movie titles and cast lists, the system uses SPARQL Construct queries to pull data directly from DBpedia (the structured version of Wikipedia).

Ontology Enrichment Process

Figure: The process of fetching ground facts from Linked Open Data (LOD).

Integrating NLP with Ontology

To make the ontology "readable" for computers processing human text, the authors developed a pipeline using the GATE (General Architecture for Text Engineering) framework.

  • Tokenization & POS Tagging: Breaking sentences into words and identifying nouns, verbs, etc.
  • Onto Root Gazetteer (ORG): This is the "bridge." It matches the root of a word in a review (e.g., "acting") to a concept in the ontology (e.g., the class Stars).
  • JAPE Rules: Hand-crafted rules that help the system understand the context of these matches and group them as "Feature Instances."

System Pipeline

Figure: The logic flow from raw review text to semantically annotated output.

Experimental Insights

By combining semantic knowledge with NLP, the approach allows for Information Retrieval that was previously impossible. Users can query the system not just for "good movies," but for specific insights like "Which movies released in 1995 have a high sentiment score for their screenplay?"

The methodology successfully:

  • Mapped various synonyms (film, show, picture) to a single concept.
  • Linked specific entities (e.g., "The Addiction 1995") to their broader metadata (Director, Running Time).
  • Associated sentiments ("amazing") directly with the features they described ("Sally Lee").

Critical Analysis & Conclusion

This work marks a significant shift from Statistical NLP to Semantic NLP. The primary strength is its Scalability: as DBpedia grows, so does the knowledge base of the opinion mining tool, without requiring more manual labeling.

Limitations: The current framework relies on hand-crafted JAPE rules for syntactic parsing, which might be rigid when dealing with the highly informal or "slang" language often found in social media reviews.

Future Outlook: The authors suggest that this SKB (Semantic Knowledge-Based) approach could be expanded to other complex domains like hospitality or consumer electronics, where feature sets are large and constantly evolving. This represents a step toward a truly "intelligent" web that understands not just the words we say, but the entities we are talking about.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize DBpedia or Wikidata for automatic feature extraction in aspect-based sentiment analysis (ABSA).
  • Which paper first introduced the Marl ontology for sentiment annotation, and how does the MOR model extend its original framework?
  • Explore studies that apply ontology-based semantic modeling for opinion mining in non-textual domains, such as audio or video reviews.
Contents
MOR Model: Bridging the Gap Between Linked Open Data and Opinion Mining
1. TL;DR
2. The "Knowledge Gap" in Sentiment Analysis
3. Methodology: The MOR Framework
3.1. Architecture and Enrichment
4. Integrating NLP with Ontology
5. Experimental Insights
6. Critical Analysis & Conclusion