Triaging Change Requests: Why Code Authorship Trumps Heavyweight Mining

Triaging incoming change requests: Bug or commit history, or code authorship?

2012-09-01
Mario Linares Vásquez, Kamal Hossen, Hoang Dang, Huzefa H. Kagdi, Malcom Gethers, Denys Poshyvanyk
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a novel approach for triaging incoming software change requests (bugs or new features) by recommending expert developers through a combination of Information Retrieval (LSI) and source code authorship analysis. It introduces a lightweight method that specifically extracts author information from source code header comments to rank implementation expertise without the need for historical repository mining.

TL;DR

Assigning the right developer to a bug report is critical for software maintenance. While most SOTA methods mine years of commit history or bug report archives, this paper proposes a "back-to-basics" approach: extracting authors directly from source code header comments. By combining Information Retrieval (IR) with simple authorship frequency, the authors achieve results that equal or outperform complex Machine Learning models and version-history miners.

Positioning: This work is a "paradigm-shifter" that challenges the necessity of heavyweight repository mining (MSR) for effective developer recommendation.

The "History" Problem: Why Tradition is Heavy

Current developer recommendation systems typically fall into two camps:

  1. The ML Camp: Training classifiers on thousands of past bug reports.
  2. The History Camp (xFinder): Mining the entire commit history to see who touched what file most often.

Both require "big data" from the past. But what if the repository history is messy, inaccessible, or non-existent? The authors argue that the most relevant expert is often the person who literally wrote the code—the person whose name is etched in the file's header comments.

Methodology: IR Meets Header Extraction

The proposed workflow is elegant in its simplicity:

  1. Concept Location: When a new bug report arrives, the system uses Latent Semantic Indexing (LSI) to find the most "conceptually similar" source files based on the bug's description.
  2. Authorship Extraction: Instead of checking who committed the file, it reads the file itself. It extracts names from the header comments (e.g., Author: mvw).
  3. Frequency Ranking: It ranks developers based on how many "relevant" files they are listed in. If there’s a tie, it looks at the rank of the file or the lexical position of the name in the header.

Overall Workflow Architecture Figure 1: Extracting authors (mvw, jaap) directly from the header comments of a Java class.

Experiments: Lightweight vs. Heavyweight

The authors tested their method against ML (SVM-based classification) and xFinder (Commit mining) across three open-source projects: ArgoUML, jEdit, and MuCommander.

Key Findings:

  • High Accuracy: In MuCommander and jEdit, the Authorship approach significantly outperformed Machine Learning in precision and recall (specifically for Top-3 to Top-10 recommendations).
  • Efficiency: The entire process took only ~90 seconds on a standard desktop—magnitudes faster than training a global SVM or parsing millions of commits.
  • Dynamic Analysis?: Interestingly, adding dynamic analysis (execution traces) to filter IR results did NOT significantly improve accuracy, suggesting the static authorship signal is already remarkably robust.

Performance Comparison Figure 2: Precision-Recall curves showing the Authorship approach (blue) competing strongly against xFinder (commit-based) and ML (bug-history-based).

Critical Insights & Conclusion

This paper serves as a reminder that Information Retrieval is often about finding the right signal, not just the biggest dataset.

Takeaways:

  • The "Gold" in Comments: Header comments are more than just legal boilerplate; they are high-precision indicators of implementation expertise.
  • Simplicity Wins: By bypassing the need for repository mining, this tool can be deployed on a single snapshot of source code, making it highly portable.

Limitations: The method's Achilles' heel is its dependency on developer discipline. If a team doesn't maintain header comments, the "author" signal disappears. However, in the vast world of Open Source (as shown in the study), this metadata is a goldmine waiting to be tapped.

Future Outlook: The next step is "Hybrid Triaging"—combining commit logs with authorship headers to see if the "Original Author" vs. "Frequent Maintainer" provides a better expertise profile than either one alone.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize source code metadata or "inline" documentation to assist in automated software maintenance tasks like developer recommendation or bug localization.
  • Which study first introduced the use of Latent Semantic Indexing (LSI) for concept location in source code, and how has this authorship-based approach evolved from that baseline?
  • Examine research that evaluates the accuracy of code ownership and authorship information in commercial versus open-source projects to see if the header-comment method remains viable in proprietary environments.
Contents
Triaging Change Requests: Why Code Authorship Trumps Heavyweight Mining
1. TL;DR
2. The "History" Problem: Why Tradition is Heavy
3. Methodology: IR Meets Header Extraction
4. Experiments: Lightweight vs. Heavyweight
4.1. Key Findings:
5. Critical Insights & Conclusion