Triaging Change Requests: Why Code Authorship Trumps Heavyweight Mining
Triaging incoming change requests: Bug or commit history, or code authorship?
The paper presents a novel approach for triaging incoming software change requests (bugs or new features) by recommending expert developers through a combination of Information Retrieval (LSI) and source code authorship analysis. It introduces a lightweight method that specifically extracts author information from source code header comments to rank implementation expertise without the need for historical repository mining.
TL;DR
Assigning the right developer to a bug report is critical for software maintenance. While most SOTA methods mine years of commit history or bug report archives, this paper proposes a "back-to-basics" approach: extracting authors directly from source code header comments. By combining Information Retrieval (IR) with simple authorship frequency, the authors achieve results that equal or outperform complex Machine Learning models and version-history miners.
Positioning: This work is a "paradigm-shifter" that challenges the necessity of heavyweight repository mining (MSR) for effective developer recommendation.
The "History" Problem: Why Tradition is Heavy
Current developer recommendation systems typically fall into two camps:
- The ML Camp: Training classifiers on thousands of past bug reports.
- The History Camp (xFinder): Mining the entire commit history to see who touched what file most often.
Both require "big data" from the past. But what if the repository history is messy, inaccessible, or non-existent? The authors argue that the most relevant expert is often the person who literally wrote the code—the person whose name is etched in the file's header comments.
Methodology: IR Meets Header Extraction
The proposed workflow is elegant in its simplicity:
- Concept Location: When a new bug report arrives, the system uses Latent Semantic Indexing (LSI) to find the most "conceptually similar" source files based on the bug's description.
- Authorship Extraction: Instead of checking who committed the file, it reads the file itself. It extracts names from the header comments (e.g.,
Author: mvw). - Frequency Ranking: It ranks developers based on how many "relevant" files they are listed in. If there’s a tie, it looks at the rank of the file or the lexical position of the name in the header.
Figure 1: Extracting authors (mvw, jaap) directly from the header comments of a Java class.
Experiments: Lightweight vs. Heavyweight
The authors tested their method against ML (SVM-based classification) and xFinder (Commit mining) across three open-source projects: ArgoUML, jEdit, and MuCommander.
Key Findings:
- High Accuracy: In MuCommander and jEdit, the Authorship approach significantly outperformed Machine Learning in precision and recall (specifically for Top-3 to Top-10 recommendations).
- Efficiency: The entire process took only ~90 seconds on a standard desktop—magnitudes faster than training a global SVM or parsing millions of commits.
- Dynamic Analysis?: Interestingly, adding dynamic analysis (execution traces) to filter IR results did NOT significantly improve accuracy, suggesting the static authorship signal is already remarkably robust.
Figure 2: Precision-Recall curves showing the Authorship approach (blue) competing strongly against xFinder (commit-based) and ML (bug-history-based).
Critical Insights & Conclusion
This paper serves as a reminder that Information Retrieval is often about finding the right signal, not just the biggest dataset.
Takeaways:
- The "Gold" in Comments: Header comments are more than just legal boilerplate; they are high-precision indicators of implementation expertise.
- Simplicity Wins: By bypassing the need for repository mining, this tool can be deployed on a single snapshot of source code, making it highly portable.
Limitations: The method's Achilles' heel is its dependency on developer discipline. If a team doesn't maintain header comments, the "author" signal disappears. However, in the vast world of Open Source (as shown in the study), this metadata is a goldmine waiting to be tapped.
Future Outlook: The next step is "Hybrid Triaging"—combining commit logs with authorship headers to see if the "Original Author" vs. "Frequent Maintainer" provides a better expertise profile than either one alone.
