Unified Feedback: Bringing Semantic Intelligence to Google Play and Twitter Bug Reports
Semantic analysis of issues on Google play and Twitter
This paper proposes a semantic analysis framework to automatically identify similarities and differences between mobile app bug reports on Google Play Store and Twitter. By leveraging Word2Vec embeddings and Jensen-Shannon Distance, the system aligns cross-platform user feedback at three granularity levels: topics, n-grams, and individual comments.
TL;DR
Developers are often overwhelmed by user feedback scattered across the Google Play Store and Twitter. This paper introduces an automated framework that uses Semantic Word Embeddings and Probability Distribution Distances to bridge the gap between these platforms, allowing developers to identify overlapping bug reports and unique issues without manual sifting.
The "Vocabulary Trap" in App Feedback
For a mobile developer, user feedback is gold. However, users on Twitter talk differently than those on the Play Store. A user might tweet "The app crashes when I take a photo," while a Play Store review says "Camera module bug in version 2.1."
Standard keyword-based systems (Bag-of-Words) often treat these as completely different issues because the words don't match. This leads to redundant work—where the same bug is tracked twice—or missed signals, where a unique issue on Twitter is buried under the noise of the App Store.
Methodology: A Triple-Layer Semantic Alignment
The authors argue that to truly understand user intent, we must look at the semantics (meaning) rather than just the syntax (words). Their framework operates at three distinct levels of granularity:
- Topics (Macro Level): Using Latent Dirichlet Allocation (LDA) to group feedback into broad categories like "Login Issues" or "UI Glitches."
- N-grams (Mid Level): Analyzing 3-word and 4-word continuous sequences (e.g., "app not working") to find specific recurring phrases.
- Comments (Micro Level): Comparing individual tweets directly to reviews.
The Tech Stack
- Embeddings: They utilize Word2Vec to transform words into high-dimensional vectors. This ensures that "Error" and "Issue" are mathematically close to each other.
- Similarity Metrics:
- Cosine Similarity is used for topics and n-grams.
- Jensen-Shannon Distance (JSD) is used for comments. JSD measures how much two probability distributions (topic mixes) overlap.
The JSD formula used to calculate the distance between a tweet's topic distribution and a review's topic distribution.
Experimental Results & Expert Validation
To prove the system works, the researchers compared the machine's "similarity scores" with scores given by human experts (software engineers with 10+ years of experience).
| Analysis Level | Spearman's | Conclusion |
|---|---|---|
| Topics | 0.57 | Strong positive correlation; highly effective at high-level grouping. |
| 4-grams | 0.38 | Moderate correlation; useful for catching specific technical phrases. |
| Comments | -0.16 (p=1.1e-03) | Statistically significant relationship between JSD distance and human similarity. |

The results confirmed that the framework effectively identifies when a tweet and a review are talking about the same thing, even if the wording differs.
Critical Insight & Future Outlook
The core strength of this paper lies in its multi-platform integration. Many previous works analyzed Google Play or Twitter in isolation. This work recognizes that the modern software ecosystem is multi-channel.
Limitations: While Word2Vec was SOTA at the time of this paper (2020), it lacks the context-awareness of modern Transformers (like BERT or GPT). Word2Vec gives the word "bank" the same vector regardless of whether it's a "river bank" or a "financial bank."
Future Impact: For practitioners, this framework paves the way for a Unified Developer Dashboard. Imagine a tool that automatically clusters 1,000 tweets and 5,000 reviews into 5 "Actionable Bug Clusters," saving hundreds of hours of triage time. The next logical step is moving toward LLM-based summarization to provide developers with a natural language summary of these cross-platform issues.
Published in ICSE '20: 42nd International Conference on Software Engineering.
