Automated Traceability: Moving from Search Results to Binary Decisions
Automating traceability link recovery through classification
This paper proposes a novel approach to Automating Traceability Link Recovery (TLR) by reframing it as a binary classification problem rather than a traditional Information Retrieval (IR) task. By leveraging machine learning models like Random Forest, the method achieves a high recall (up to 92.7%) in identifying valid links between software artifacts.
TL;DR
Software Traceability Link Recovery (TLR) is traditionally treated as an Information Retrieval (IR) problem, leaving developers to sift through "top-K" results. This paper proposes a paradigm shift: treating TLR as a binary classification task. By combining similarity scores with document statistics and Query Quality (QQ) metrics, the proposed model can automatically label links as valid or invalid, achieving up to 92.7% recall and significantly reducing manual verification effort.
The Bottleneck of Manual Verification
In the software lifecycle, linking requirements to source code or test cases (TLR) is vital for impact analysis. However, current SOTA IR methods (like VSM or LSI) provide ranked lists. The problem? The stakeholder is still the bottleneck. A human must inspect every suggested link to prune false positives. When datasets scale, the manual effort required to find the "needle in the haystack" makes traceability unsustainable for large-scale agile projects.
Methodology: The Classifier's Lens
The core insight of this work is that a link's validity isn't just about textual similarity; it's about the context of the artifacts. The author extracts three types of features to feed into machine learning classifiers:
- IR Ranking: Scores from VSM, BM25, and language smoothing models (Jelinek Mercer/Dirichlet).
- Query Quality (QQ): Features that determine if an artifact is a "bad query." If two documents are dissimilar, is it because the link is invalid, or because the requirement was poorly written?
- Document Statistics: Basic metrics like vocabulary size and term overlap percentages.

Tackling Class Imbalance
Valid links are rare (averaging only ~8% of all possible pairs). To prevent the model from simply predicting "invalid" for everything, the author tests two strategies:
- SMOTE: Generating synthetic examples of valid links.
- Undersampling: Reducing the number of invalid link examples in the training set.
Critical Results & Performance
The author evaluated several algorithms, including J48, Naive Bayes, and Random Forest.

Key Insights from the Benchmarks:
- Random Forest is the winner: It consistently showed the best balance between finding valid links (TPR) and avoiding false alarms (FPR).
- The Precision-Recall Trade-off: Using SMOTE led to an incredibly low FPR (0.017), meaning the links it did find were almost certainly correct. Undersampling caught more links (92.7% recall) but introduced more noise (12.2% FPR).
Critical Analysis & Conclusion
This work is a significant step toward "Zero-Touch Traceability." By reformulating the problem, it moves the software engineering community closer to tools that don't just suggest links but actually maintain them.
Limitations:
- Dependency on Labeled Data: Classification requires ground truth. For a brand-new project with no history, a "Cold Start" problem exists.
- Feature Engineering: The current features are relatively "shallow" (word counts, basic IR). Incorporating semantic embeddings (like CodeBERT) could likely drive the FPR even lower.
Takeaway: The future of software engineering tools lies in automation that removes the human from the loop of trivial verification. This paper provides the foundational framework for that transition in the realm of traceability.
