ENRL: Solving the "When to Stop" Dilemma in Software Traceability Recovery
Estimating the number of remaining links in traceability recovery
The paper introduces ENRL (Estimation of the Number of Remaining Links), a machine learning-based approach to estimate how many positive traceability links remain in a ranked list generated by NLP techniques. By training classifiers on a fraction of manually inspected links, it achieves high accuracy across industrial and academic datasets, providing a stopping criterion for requirements analysts.
TL;DR
Automated traceability recovery often leaves analysts staring at thousands of candidate links with no idea when they've found "enough" of them. This paper introduces ENRL, a supervised machine learning framework that predicts the number of remaining positive links in a ranked list. By learning from the analyst's first few manual inspections, ENRL provides an objective stopping rule that can achieve a Mean Relative Error (MRE) of less than 0.15, significantly reducing manual overhead in large-scale industrial projects.
Problem & Motivation: The Infinite List Problem
Traceability links (e.g., connecting requirements to source code) are the backbone of software maintenance and safety compliance. However, for a system with just 2,500 requirements, there are over 3 million possible pairs to inspect.
Current SOTA Natural Language Processing (NLP) tools can rank these pairs by similarity, but they are never 100% precise. Analysts are forced to make a blind tradeoff:
- Stop too early: Miss critical links, leading to duplicate code or failed safety audits.
- Stop too late: Waste hours inspecting "noise" (false positives) at the bottom of the list.
The authors observed that while we have tools to find links, we lack tools to quantify what's left.
Methodology: The ENRL Framework
The core insight of ENRL (Estimation of the Number of Remaining Links) is that the human analyst's initial manual work is a goldmine for training data.
The 6-Step Workflow
- Extraction: Generate all possible artifact pairs.
- NLP Ranking: Use techniques like VSM (Vector Space Model) or LSA (Latent Semantic Analysis) to score pairs.
- Manual Inspection: The analyst verifies the top links.
- ML Training: An ML classifier is trained using the similarity scores as features and the analyst's "Positive/Negative" labels as targets.
- Prediction: The trained model classifies the remaining (unseen) pairs.
- Estimation: The sum of predicted "Positive" links becomes the stopping guide.

Experiments & Results
The researchers tested ENRL on three major datasets, including one from the aerospace industry (Selex SI). They compared 7 different ML classifiers (like LogitBoost, IBk, and NaiveBayes) and 12 NLP configurations.
Key Performance metrics:
- High Accuracy: For the industrial Selex SI dataset, the FLR (Fuzzy Lattice Reasoning) classifier achieved an MRE as low as 0.03.
- Minimal Data Requirement: In some cases (e.g., the EasyClinic dataset), the model became accurate after the analyst inspected only 20% of the list.
- The "Univariate" Surprise: Surprisingly, multivariate models (combining all NLP techniques) were less accurate than univariate models using a single well-chosen NLP technique. The authors suggest that adding too many NLP features introduces noise that confuses the classifiers.
Figure: As shown in the charts above, accuracy (MRE) often stabilizes or peaks quickly, suggesting that analysts don't need to finish the whole list to get a reliable estimate.
SOTA Comparison
Unlike traditional Capture-Recapture models used in biology or defect estimation, which require multiple independent analysts to work simultaneously, ENRL works with a single analyst following a ranked list—a much more realistic scenario for industrial software engineering.
Critical Analysis & Conclusion
Takeaways
- Classifier Selection Matters: No single ML algorithm won every time. IBk (K-nearest neighbor) excelled in academic code-based datasets, while FLR was superior for formal industrial requirements.
- Correlation is Key: The "Estimated MRE" (measured via cross-validation on the known links) was highly correlated with the "Actual MRE." This means analysts can trust the tool's internal confidence score to decide if the estimate is reliable.
Limitations
The study used "out-of-the-box" versions of Weka's classifiers. Tuning these hyperparameters specifically for traceability data would likely yield even higher accuracy. Additionally, the performance of the model can fluctuate if the ranked list contains "bursts" of difficult links that the NLP tool failed to score correctly.
Impact
ENRL transforms traceability from a never-ending manual chore into a data-driven process. It gives project managers a "Gas Gauge" for their traceability efforts, allowing them to say with statistical confidence: "We have found 95% of the links; we can stop now."
