Beyond the Explicit: Boosting Opinion Retrieval via Coreference Resolution and MBL

Improving opinion retrieval in social media by combining features-based coreferencing and memory-based learning q

2014-12-16
John Atkinson, Gonzalo Salas, Alejandro Figueroa
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a novel adaptive approach for opinion retrieval in social media by combining feature-based linguistic coreferencing with Memory-Based Learning (MBL). It specifically targets the identification of implicit entities and features in short, informal texts like Twitter messages, achieving a Mean Average Precision (MAP) of 0.532, significantly outperforming traditional supervised classifiers and lexicon-based methods.

TL;DR

When you search for opinions about "Bank X" on Twitter, you might miss half the relevant data because users often use "it," "they," or "this institution" instead of the brand name. This paper introduces an adaptive framework that combines Memory-Based Learning (MBL) with linguistic coreferencing to track these implicit mentions across social media threads, boosting retrieval precision by nearly 18% over traditional baselines.

The Problem: The "Implicit" Gap in Social Media

Traditional IR (Information Retrieval) models are literal. If a tweet says "The new phone is great, but its battery dies fast," a search for "battery" works, but a search for "phone" might struggle to link the battery complaint specifically to the object if the discourse structure is complex. In social media, this is exacerbated by:

  • Short message constraints: Users omit names to save space.
  • Threaded nature: A reply might just say "True, I hate it," where "it" refers to a topic mentioned three levels up in the thread.
  • Noisy syntax: Use of @, #, and RT breaks standard NLP parsers.

The authors identify that nearly 14% of key objects in opinions are expressed as pronoun anaphoras, creating a massive data leak for sentiment analysis tools.

Methodology: Connecting the Dots

The researchers propose a three-stage pipeline (Message Retrieval -> Preprocessing -> Referencing Analysis) that treats coreference resolution as a classification task.

1. Harnessing Thread Hierarchies

Unlike previous methods that look at tweets in isolation, this approach builds a hierarchy of messages. By treating sub-comments as "Reply Messages" and top-level posts as "Original Messages," the system uses the conversation tree as a roadmap for finding antecedents.

2. Feature-Rich Representation

The core of the system is the Memory-Based Learning (MBL) classifier. It doesn't just look at word frequency; it evaluates 15 distinct linguistic features:

  • Lexical: String matches and "Head Match" (do they share the same nucleus?).
  • Grammatical: Gender and number agreement (though the paper notes users often break these rules in informal settings).
  • Semantic: Classifying entities into categories like "Person" or "Organization" via LabeledLDA.

Model Architecture Placeholder Fig 1: The proposed sentiment-aware retrieval architecture integrating coreference resolution.

Experiments and Insights

The study compared MBL against SVMs and traditional retrieval methods (Lexicon-based and Probabilistic).

Key Result: MBL vs. SVM

MBL significantly outperformed SVMs in both time efficiency and accuracy. While the SVM took 240 seconds to process, MBL finished in 67 seconds with superior Precision (0.92 vs 0.83 for NP antecedents).

Performance in Opinion Retrieval

The real-world test involved 50 queries across domains like politics and electronics.

MethodMAP
Our Linguistic MBL-based Method0.532
Supervised Text Classifier0.448
Lexicon-based Approach0.453
Probabilistic Approach0.429

The "Head Match" feature (Feature 7) was found to be the single most influential factor. If the central "nucleus" of two phrases matched, the probability of them being coreferent skyrocketed.

Result Analysis Fig 2: Performance metrics showing the impact of incremental feature addition.

Critical Analysis & Conclusion

This paper makes a strong case that discourse structure matters. By shifting focus from individual keywords to "referencing chains," the authors solve a major pain point in social media mining.

Limitations:

  • Informal Grammar: The study found that feature 12 (number agreement) often fails on Twitter because users might refer to a singular "company" as "they" (e.g., "The bank is slow, they need to hire more people").
  • Normalisation Dependency: The performance drops by ~8% on raw, unnormalized tweets, showing that the method still relies heavily on clean text.

Future Outlook: As we move into the era of LLMs, the "memory-based" logic of looking at similar past instances remains relevant, but the manual feature engineering (15 specific rules) could likely be replaced by attention mechanisms. However, the use of thread hierarchies as a structural constraint remains a "best practice" for anyone working with social media data today.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning or Transformer-based models for coreference resolution specifically in microblogging and short-text environments.
  • Which study first established Memory-Based Learning (MBL) as a viable alternative for NLP tasks like POS tagging or chunking, and how does this paper adapt those principles for discourse analysis?
  • Explore how contemporary Large Language Models (LLMs) handle the "implicit entity" problem in social media compared to the feature-engineered approach presented in this paper.
Contents
Beyond the Explicit: Boosting Opinion Retrieval via Coreference Resolution and MBL
1. TL;DR
2. The Problem: The "Implicit" Gap in Social Media
3. Methodology: Connecting the Dots
3.1. 1. Harnessing Thread Hierarchies
3.2. 2. Feature-Rich Representation
4. Experiments and Insights
4.1. Key Result: MBL vs. SVM
4.2. Performance in Opinion Retrieval
5. Critical Analysis & Conclusion