Beyond Rankings: Automating Traceability Link Validation with Machine Learning

A Machine Learning Approach for Determining the Validity of Traceability Links

2017-05-01
Chris Mills, Sonia Haiduc
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a supervised machine learning framework for Traceability Link Recovery (TLR) that automatically classifies links as valid or invalid. By combining traditional Text Retrieval (TR) rankings with Query Quality (QQ) metrics, the approach achieves a high average classification accuracy of 95% across three software systems.

TL;DR

In software engineering, maintaining "links" between requirements and code is vital but labor-intensive. This paper moves the needle from ranking potential links to automatically classifying them. By combining bi-directional search rankings with Query Quality (QQ) metrics, the researchers achieved a 95% accuracy in identifying valid links, significantly reducing the manual burden on developers.

The Problem: The "List Fatigue" in Traceability

Traceability Link Recovery (TLR) is essential for impact analysis and compliance. However, the status quo is frustrating:

  1. Manual Overload: Standard Information Retrieval (IR) tools give developers a long list of "suggestions." A human must still sift through these to find the needle in the haystack.
  2. The Quality Gap: If a requirement is poorly worded (low query quality), a standard IR tool will fail to rank the relevant source code correctly, leading to missed links.

Methodology: The Hybrid Feature Approach

The authors propose a Random Forest classification model that doesn't just look at how similar two documents are, but also how "good" those documents are as sources of information.

1. Bi-directional TR Rankings

Most systems search from Requirements Code. This paper performs the search in both directions:

  • Feature A: Rank of Code Class when Requirement is the query.
  • Feature B: Rank of Requirement when Code Class is the query. This provides a "mutual confirmation" signal that strengthens the evidence for a link.

2. Query Quality (QQ) Metrics

This is the "secret sauce" of the paper. They used 14 specific metrics (applied to both sides, resulting in 28 features) to assess whether an artifact is actually a good "query." If a requirement is vague or uses non-standard terms, the model learns to be skeptical of its IR ranking.

3. Solving the Imbalance (SMOTE)

In a typical system, 95% of possible links are invalid. To prevent the model from simply guessing "invalid" every time, the authors used SMOTE (Synthetic Minority Oversampling Technique) to synthetically balance the training data, ensuring the model actually learns the characteristics of valid links.

System Data Distribution Table 1: The massive imbalance between potential and valid links in eAnci, eTour, and SMOS.

Experiments & Results

The researchers compared a model using only TR rankings against their full model (TR + QQ).

  • TR-Only Model: Achieved an average accuracy of 84.7%. While decent, it suffered from a high number of false positives (Type-I errors).
  • Full Model (TR + QQ): Accuracy jumped to 94.7%.

Enhanced Performance with QQ Metrics Table 3: Accuracy and error rates using the combined feature set. Notice the significant drop in both Type-I and Type-II errors.

Insights & Takeaways

The core insight here is that textual similarity is not enough. By accounting for the intrinsic quality of the artifacts being linked, we can filter out the noise that plagues traditional retrieval systems.

Limitations: While 95% accuracy is impressive, the evaluation was performed on relatively small datasets (eAnci, eTour, SMOS). In a massive industrial codebase with millions of potential link combinations, even a 5% error rate could lead to substantial manual cleanup.

Future Outlook: The logical next step is combining this classification approach with Large Language Models (LLMs). While LLMs are great at understanding semantics, they are computationally expensive; a Random Forest model using these lightweight features could serve as a highly efficient "first-pass" filter in professional IDEs.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning or Large Language Models (LLMs) to perform binary classification of traceability links instead of traditional Text Retrieval.
  • Which 14 specific Query Quality (QQ) metrics were originally proposed in the Haiduc et al. 2012 study, and how are they mathematically defined for software artifacts?
  • Investigate how the "bi-directional ranking" feature proposed in this paper compares to "cross-encoder" architectures in modern Natural Language Processing for document matching.
Contents
Beyond Rankings: Automating Traceability Link Validation with Machine Learning
1. TL;DR
2. The Problem: The "List Fatigue" in Traceability
3. Methodology: The Hybrid Feature Approach
3.1. 1. Bi-directional TR Rankings
3.2. 2. Query Quality (QQ) Metrics
3.3. 3. Solving the Imbalance (SMOTE)
4. Experiments & Results
5. Insights & Takeaways