LRCA: Empowering Bug Report Summarization through Crowdsourced Insights
Toward Better Summarizing Bug Reports With Crowdsourcing Elicited Attributes
This paper introduces Crowd-Attribute (CA), a novel crowdsourcing-based method for constructing effective sentence attributes to improve Bug Report Summarization. By leveraging crowd-generated reasons for sentence selection, the authors developed LRCA (Logistic Regression with Crowdsourced Attributes), which achieves SOTA performance on both the public SDS dataset and a massive new benchmark of 105,177 reports.
TL;DR
Software maintenance is a race against time, with developers often overwhelmed by massive bug reports. While automatic summarization exists, it suffers from "feature anemia"—relying on generic attributes. This paper introduces Crowd-Attribute (CA), a methodology that utilizes the collective intelligence of the crowd to "discover" 11 high-impact attributes (like code snippet detection and reporter-specific importance). The resulting model, LRCA, sets a new benchmark for accuracy across 100,000+ bug reports.
The Bottleneck of Feature Engineering
Supervised learning is only as good as the data representation (attributes) it receives. In Mining Software Repositories (MSR), researchers typically construct attributes based on personal experience or by "borrowing" them from other fields (Knowledge Transfer).
The authors' survey revealed a startling truth: most MSR papers rely on just 1-3 individuals to define these attributes. This "expert silos" approach limits the diversity of features and often misses the "hidden factors" that make a sentence in a bug report truly important.
Methodology: Mining the "Why"
The core innovation isn't just the summarization model, but the Crowd-Attribute (CA) pipeline. Instead of asking volunteers to just pick sentences, they asked "Why did you pick this?"
The HCR Framework
The researchers used a tool called CSEP to gather 332 reasons from volunteers. These reasons were processed through Heuristic Construction Rules (HCR):
- Extraction: Turning a reason like "the sentence is long" into a candidate attribute "Length".
- Filtering: Removing meaningless or redundant groups.
- Calculation/Merging: Assigning mathematical metrics (e.g., VSM similarity for "Relatedness").
Figure: The Roadmap of Crowd-Attribute and its integration into the summarization pipeline.
The 11 New Attributes
By analyzing crowd intuitions, they identified features that generic models missed:
- CODE: Identifying if a sentence is a snippet (highly discriminative).
- SWT/SWD: Similarity with the overall Topic vs. the specific Description.
- REP: Whether the sentence was written by the original reporter (usually contains more "intent").
Experimental Results: Dominating the Baseline
The authors didn't just test on the tiny public SDS dataset (36 reports); they built a massive benchmark, BRSB, with over 105,000 reports.
Performance Gains
LRCA (Logistic Regression with Crowdsourced Attributes) crushed the previous SOTA, BRC, across the board.
- Recall Improvement: +10.11% on SDS.
- Large Scale: On the BRSB datasets, LRCA achieved a significantly higher HitRate than top-performing unsupervised methods (Centroid) and BRC.
Figure: Comparison of LRCA against supervised and unsupervised baselines.
The Power of Numbers
Interestingly, the study found that the model's performance stabilizes once about 13 volunteers are involved (the "diminishing returns" of the crowd), proving that you don't need thousands of people to build a professional-grade attribute set.
Critical Analysis & Conclusion
The value of this paper lies in its movement away from "black-box" feature engineering toward a structured, human-in-the-loop (HITL) approach.
Takeaway: The success of LRCA demonstrates that human intuition—specifically the reasons why we find information useful—can be effectively encoded into statistical models.
Limitations: Despite the success, the HCR process still requires a "Requester" (expert) to interpret the crowd's reasons. In the age of LLMs, one might wonder: Could a Large Language Model replace the Requester or even the Crowd to define these attributes autonomously?
That is the next frontier for automated software engineering.
