Unified Feedback: Bringing Semantic Intelligence to Google Play and Twitter Bug Reports

Semantic analysis of issues on Google play and Twitter

2020-06-27
Aman Yadav, Fatemeh Hendijani Fard
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a semantic analysis framework to automatically identify similarities and differences between mobile app bug reports on Google Play Store and Twitter. By leveraging Word2Vec embeddings and Jensen-Shannon Distance, the system aligns cross-platform user feedback at three granularity levels: topics, n-grams, and individual comments.

TL;DR

Developers are often overwhelmed by user feedback scattered across the Google Play Store and Twitter. This paper introduces an automated framework that uses Semantic Word Embeddings and Probability Distribution Distances to bridge the gap between these platforms, allowing developers to identify overlapping bug reports and unique issues without manual sifting.

The "Vocabulary Trap" in App Feedback

For a mobile developer, user feedback is gold. However, users on Twitter talk differently than those on the Play Store. A user might tweet "The app crashes when I take a photo," while a Play Store review says "Camera module bug in version 2.1."

Standard keyword-based systems (Bag-of-Words) often treat these as completely different issues because the words don't match. This leads to redundant work—where the same bug is tracked twice—or missed signals, where a unique issue on Twitter is buried under the noise of the App Store.

Methodology: A Triple-Layer Semantic Alignment

The authors argue that to truly understand user intent, we must look at the semantics (meaning) rather than just the syntax (words). Their framework operates at three distinct levels of granularity:

  1. Topics (Macro Level): Using Latent Dirichlet Allocation (LDA) to group feedback into broad categories like "Login Issues" or "UI Glitches."
  2. N-grams (Mid Level): Analyzing 3-word and 4-word continuous sequences (e.g., "app not working") to find specific recurring phrases.
  3. Comments (Micro Level): Comparing individual tweets directly to reviews.

The Tech Stack

  • Embeddings: They utilize Word2Vec to transform words into high-dimensional vectors. This ensures that "Error" and "Issue" are mathematically close to each other.
  • Similarity Metrics:
    • Cosine Similarity is used for topics and n-grams.
    • Jensen-Shannon Distance (JSD) is used for comments. JSD measures how much two probability distributions (topic mixes) overlap.

Framework Logic - Formula The JSD formula used to calculate the distance between a tweet's topic distribution and a review's topic distribution.

Experimental Results & Expert Validation

To prove the system works, the researchers compared the machine's "similarity scores" with scores given by human experts (software engineers with 10+ years of experience).

Analysis LevelSpearman's Conclusion
Topics0.57Strong positive correlation; highly effective at high-level grouping.
4-grams0.38Moderate correlation; useful for catching specific technical phrases.
Comments-0.16 (p=1.1e-03)Statistically significant relationship between JSD distance and human similarity.

Experimental Correlation Table

The results confirmed that the framework effectively identifies when a tweet and a review are talking about the same thing, even if the wording differs.

Critical Insight & Future Outlook

The core strength of this paper lies in its multi-platform integration. Many previous works analyzed Google Play or Twitter in isolation. This work recognizes that the modern software ecosystem is multi-channel.

Limitations: While Word2Vec was SOTA at the time of this paper (2020), it lacks the context-awareness of modern Transformers (like BERT or GPT). Word2Vec gives the word "bank" the same vector regardless of whether it's a "river bank" or a "financial bank."

Future Impact: For practitioners, this framework paves the way for a Unified Developer Dashboard. Imagine a tool that automatically clusters 1,000 tweets and 5,000 reviews into 5 "Actionable Bug Clusters," saving hundreds of hours of triage time. The next logical step is moving toward LLM-based summarization to provide developers with a natural language summary of these cross-platform issues.


Published in ICSE '20: 42nd International Conference on Software Engineering.

Find Similar Papers

Try Our Examples

  • Find recent studies that use Large Language Models (LLMs) instead of Word2Vec for cross-platform app review classification and sentiment analysis.
  • Which paper first introduced the use of Jensen-Shannon Distance for measuring document similarity in software engineering repositories?
  • Explore how automated feedback aggregation frameworks like this are being integrated into DevOps pipelines for continuous software evolution.
Contents
Unified Feedback: Bringing Semantic Intelligence to Google Play and Twitter Bug Reports
1. TL;DR
2. The "Vocabulary Trap" in App Feedback
3. Methodology: A Triple-Layer Semantic Alignment
3.1. The Tech Stack
4. Experimental Results & Expert Validation
5. Critical Insight & Future Outlook