ENRL: Solving the "When to Stop" Dilemma in Software Traceability Recovery

Estimating the number of remaining links in traceability recovery

2016-10-20
Davide Falessi, Massimiliano Di Penta, Gerardo Canfora, Giovanni Cantone
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces ENRL (Estimation of the Number of Remaining Links), a machine learning-based approach to estimate how many positive traceability links remain in a ranked list generated by NLP techniques. By training classifiers on a fraction of manually inspected links, it achieves high accuracy across industrial and academic datasets, providing a stopping criterion for requirements analysts.

TL;DR

Automated traceability recovery often leaves analysts staring at thousands of candidate links with no idea when they've found "enough" of them. This paper introduces ENRL, a supervised machine learning framework that predicts the number of remaining positive links in a ranked list. By learning from the analyst's first few manual inspections, ENRL provides an objective stopping rule that can achieve a Mean Relative Error (MRE) of less than 0.15, significantly reducing manual overhead in large-scale industrial projects.

Problem & Motivation: The Infinite List Problem

Traceability links (e.g., connecting requirements to source code) are the backbone of software maintenance and safety compliance. However, for a system with just 2,500 requirements, there are over 3 million possible pairs to inspect.

Current SOTA Natural Language Processing (NLP) tools can rank these pairs by similarity, but they are never 100% precise. Analysts are forced to make a blind tradeoff:

  1. Stop too early: Miss critical links, leading to duplicate code or failed safety audits.
  2. Stop too late: Waste hours inspecting "noise" (false positives) at the bottom of the list.

The authors observed that while we have tools to find links, we lack tools to quantify what's left.

Methodology: The ENRL Framework

The core insight of ENRL (Estimation of the Number of Remaining Links) is that the human analyst's initial manual work is a goldmine for training data.

The 6-Step Workflow

  1. Extraction: Generate all possible artifact pairs.
  2. NLP Ranking: Use techniques like VSM (Vector Space Model) or LSA (Latent Semantic Analysis) to score pairs.
  3. Manual Inspection: The analyst verifies the top links.
  4. ML Training: An ML classifier is trained using the similarity scores as features and the analyst's "Positive/Negative" labels as targets.
  5. Prediction: The trained model classifies the remaining (unseen) pairs.
  6. Estimation: The sum of predicted "Positive" links becomes the stopping guide.

ENRL Process Architecture

Experiments & Results

The researchers tested ENRL on three major datasets, including one from the aerospace industry (Selex SI). They compared 7 different ML classifiers (like LogitBoost, IBk, and NaiveBayes) and 12 NLP configurations.

Key Performance metrics:

  • High Accuracy: For the industrial Selex SI dataset, the FLR (Fuzzy Lattice Reasoning) classifier achieved an MRE as low as 0.03.
  • Minimal Data Requirement: In some cases (e.g., the EasyClinic dataset), the model became accurate after the analyst inspected only 20% of the list.
  • The "Univariate" Surprise: Surprisingly, multivariate models (combining all NLP techniques) were less accurate than univariate models using a single well-chosen NLP technique. The authors suggest that adding too many NLP features introduces noise that confuses the classifiers.

Accuracy vs Training Set Size Figure: As shown in the charts above, accuracy (MRE) often stabilizes or peaks quickly, suggesting that analysts don't need to finish the whole list to get a reliable estimate.

SOTA Comparison

Unlike traditional Capture-Recapture models used in biology or defect estimation, which require multiple independent analysts to work simultaneously, ENRL works with a single analyst following a ranked list—a much more realistic scenario for industrial software engineering.

Critical Analysis & Conclusion

Takeaways

  • Classifier Selection Matters: No single ML algorithm won every time. IBk (K-nearest neighbor) excelled in academic code-based datasets, while FLR was superior for formal industrial requirements.
  • Correlation is Key: The "Estimated MRE" (measured via cross-validation on the known links) was highly correlated with the "Actual MRE." This means analysts can trust the tool's internal confidence score to decide if the estimate is reliable.

Limitations

The study used "out-of-the-box" versions of Weka's classifiers. Tuning these hyperparameters specifically for traceability data would likely yield even higher accuracy. Additionally, the performance of the model can fluctuate if the ranked list contains "bursts" of difficult links that the NLP tool failed to score correctly.

Impact

ENRL transforms traceability from a never-ending manual chore into a data-driven process. It gives project managers a "Gas Gauge" for their traceability efforts, allowing them to say with statistical confidence: "We have found 95% of the links; we can stop now."

Find Similar Papers

Try Our Examples

  • Find recent papers (post-2016) that utilize deep learning or transformer-based embeddings (like BERT) to improve the Estimation of the Number of Remaining Links (ENRL) in traceability recovery.
  • Which study first introduced the use of "capture-recapture" models for software defect estimation, and how do those assumptions contrast with modern supervised ML approaches in traceability?
  • Explore if the ENRL methodology has been successfully applied to other software engineering ranking tasks, such as bug localization or duplicate bug report detection.
Contents
ENRL: Solving the "When to Stop" Dilemma in Software Traceability Recovery
1. TL;DR
2. Problem & Motivation: The Infinite List Problem
3. Methodology: The ENRL Framework
3.1. The 6-Step Workflow
4. Experiments & Results
4.1. Key Performance metrics:
4.2. SOTA Comparison
5. Critical Analysis & Conclusion
5.1. Takeaways
5.2. Limitations
5.3. Impact