F007: Pinpointing Recurring Faults in 20 Million Lines of Code with Machine Learning

12477_Identifying Recurring Faulty Functions in Field Traces of a Large Industrial Software System.

Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces F007 (Faulty Function Finder), a supervised machine learning approach using C4.5 Decision Trees to identify recurring faulty functions in field failure traces. Evaluated on a massive industrial system with 20 million lines of code (LOC), it achieves a diagnostic accuracy of 90% to 97% for recurring faults.

TL;DR

Legacy software systems often suffer from "Old Wine in New Bottles"—failure reports that look new but are actually caused by previously resolved faults. F007 is a diagnostic tool that uses Decision Trees to map incoming field failure traces to known faulty functions. Tested on a massive IBM-scale system (20M LOC), it identifies recurring bugs with up to 97% accuracy, allowing developers to skip the manual "detective work" and jump straight to the fix.

The Motivation: Why are we still fixing the same bugs?

In a perfect world, once a bug is fixed, it stays fixed. In the world of industrial enterprise software, 50% to 90% of field failures are recurring. Users don't update their systems, vendors take time to push patches, and the same faulty function crashes systems across the globe for years.

The problem isn't that we don't know the fix; it's that we can't recognize the bug. Typical execution traces are massive (multi-gigabyte files), and manual analysis of function calls is a needle-in-a-haystack problem. Prior clustering-based approaches only grouped similar traces—they didn't tell you exactly which function to fix.

Methodology: Reducing Complexity to Essence

The researchers behind F007 (Faulty Function Finder) realized that execution sequences (Path A -> Path B) are computationally expensive to track but don't necessarily provide more diagnostic value than event frequency.

1. Feature Engineering

Instead of mapping the entire execution path, F007 transforms a trace into a probability vector of individual events:

  • Function Entry/Exit: Who was called?
  • Probe Points: Specific monitoring hooks.
  • Error Codes: Exceptions thrown.

Methodology Workflow

2. The One-Against-All Classifier

Since a trace can have multiple labels (multiple faulty functions), F007 uses a one-against-all approach. If you have 100 known faulty functions, you train 100 binary C4.5 Decision Trees. Each tree asks: "Does this trace look like the specific signature of Function X?"

3. Radical Data Reduction

One of the most surprising findings was Heuristic C: By focusing only on functions with a high standard deviation (standard deviation > 400), the researchers could discard 90% of the trace data while maintaining accuracy. This is a game-changer for systems that can't afford to store terabytes of diagnostic logs.

Experiments: Real-World Battle Testing

The system was tested on a proprietary commercial application with:

  • 20 Million+ LOC
  • 200,000+ Functions
  • 1 Million+ Users

Key Results

  • Precision: By reviewing only the top 8 recommended functions (out of 200,000!), developers found the bug in ~90% of recurring cases.
  • Cross-Release Success: Models trained on Release 1 could identify the same faulty functions in Release 2 and 3 with up to 97% accuracy. This proves that "faulty genes" persist across versions.

Experimental Results Comparison

The graph highlights that F007 significantly outperforms the "Straw Man" (random ranking) and traditional clustering, which often groups many unrelated functions together.

Critical Analysis & Deep Insights

Why does it work?

The physical intuition here is that a fault—especially a non-crashing one—creates a specific behavioral signature in the system's execution profile. Even if the exact code path varies slightly, the statistical distribution of function calls remains a high-fidelity "fingerprint" of the defect.

Limitations

  • The "Cold Start" Problem: F007 cannot identify new faults. It is a tool for recognizing old enemies, not discovering new ones.
  • Data Imbalance: If one bug has 1,000 traces and another has 5, the model will naturally be biased toward the common bug.

Conclusion and Future Outlook

F007 demonstrates that machine learning doesn't need to be complex to be effective at an industrial scale. By treating function calls as a "bag of words" problem (similar to NLP), the researchers solved a massive software engineering bottleneck.

Future iterations will likely integrate online learning, allowing the model to update itself every time a developer closes a new bug report, making the "Faulty Function Finder" a living, breathing part of the CI/CD pipeline.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply supervised machine learning models other than Decision Trees (e.g., Deep Learning) to the problem of software fault localization in large-scale industrial traces.
  • Which 2003 paper by Podgurski et al. established the foundation for automated software failure report classification, and how does F007's feature selection specifically address its limitations?
  • Explore research that applies the F007 methodology or similar function-frequency based diagnostics to multi-modal system monitoring, including cloud infrastructure logs or microservice traces.
Contents
F007: Pinpointing Recurring Faults in 20 Million Lines of Code with Machine Learning
1. TL;DR
2. The Motivation: Why are we still fixing the same bugs?
3. Methodology: Reducing Complexity to Essence
3.1. 1. Feature Engineering
3.2. 2. The One-Against-All Classifier
3.3. 3. Radical Data Reduction
4. Experiments: Real-World Battle Testing
4.1. Key Results
5. Critical Analysis & Deep Insights
5.1. Why does it work?
5.2. Limitations
6. Conclusion and Future Outlook