Beyond Sentiment: Linking High-Quality Civic Feedback to Political Discourse
What's Public Feedback? Linking High Quality Feedback to Social Issues Using Social Media
This paper introduces a novel framework for linking high-quality public feedback from social media to specific issues discussed in political speeches. By combining supervised learning for quality assessment with a semi-supervised topic model (SLDA) for issue alignment, the authors achieve SOTA performance in ranking relevant, well-justified public opinions.
TL;DR
Government officials often struggle to process the "firehose" of social media comments following major policy speeches. This paper presents a framework that filters out "noise" to find high-quality, justified public feedback and automatically links it to specific issues (like Healthcare or Economy) using a specialized topic model called SLDA.
Background: The Web 2.0 Policy Dilemma
In the era of Web 1.0, governments spoke, and citizens listened. In Web 2.0, the surge of online forums and social networks has created a goldmine of public opinion. However, for a policy maker reading 550 comments on a National Day Rally speech, the challenge isn't finding opinions—it's finding substance. Most comments are either repetitive ("I agree!") or lack justification. This paper addresses the "Needle in the Haystack" problem: Identifying the feedback that actually provides reasoning and linking it to the correct part of the policy speech.
The Core Insight: Quality + Relevance
The researchers argue that a comment is only useful to a policy maker if it satisfies two conditions:
- High Quality: It provides justification through elaboration, comparison, or examples.
- Relevance: It clearly maps to a specific issue discussed in the speech.
1. The Quality Sieve
Instead of just looking at keywords, the authors use Discourse Features. They look for structural markers of reasoning (e.g., "because," "for example," "however"). By training a Logistic Regression model on features like verb phrase density and discourse relations, they can automatically flag comments that offer "well-justified" perspectives.
2. SLDA: Structuring the Conversation
Standard Topic Models (like LDA) often miss the mark because they don't know the "context" of the speech. The authors introduced Speech-aligned LDA (SLDA).

As shown in the graphical model above, SLDA treats each pre-defined issue in the speech as a "fixed anchor" for topics, then forces the model to align user comments to these specific anchors. This ensures that the "relevance" score is grounded in the actual content of the political address.
Experimental Battleground: Obama vs. Lee Hsien Loong
The model was tested on two diverse datasets: US President Obama’s State of the Union Speech and the Singapore Prime Minister’s National Day Rally Speech.
Key Performance Metrics
The results demonstrated that the proposed SLDA model, especially when weighted with the Quality score (), significantly boosted the precision of retrieved feedback.

- The Precision Boost: For popular issues like "Immigration," the model achieved a staggering 100% Precision@10.
- The Robustness: Unlike Bag-of-Words (BOW) models, SLDA was able to link comments containing informal language or slang to the correct formal policy topics by understanding the underlying distribution of words.
Deep Insight: Discovering "Feedback Words"
One of the most valuable outputs of this research is the extraction of "feedback words"—terms used by the public that weren't in the original speech but are highly correlated with specific issues.
- Example: In Obama's speech regarding "Taxes," the model identified the word "food" as a high-probability feedback word. This revealed that the public's primary concern regarding tax policy was its impact on food security for the poor—a nuance that a simple keyword search might miss.
Critical Analysis & Conclusion
This paper serves as a bridge between traditional Political Science (surveys) and modern Data Science. By focusing on justification rather than just sentiment, it elevates social media mining from "counting likes" to "understanding why."
Limitations: The approach currently relies on manual segmentation of the target speech. Future iterations could benefit from automated speech segmentation and the use of Transformer-based embeddings (like BERT or GPT) to better handle the nuances of informal internet slang.
Future Outlook: As e-governance portals become the primary touchpoint for civic engagement, systems like SLDA will be essential for turning millions of scattered comments into a coherent "Public Report Card" that governments can actually act upon.
