BUDGET: Automating the Ground Truth for Architecture Traceability Research
BUDGET: A Tool for Supporting Software Architecture Traceability Research
BUDGET is a web-based tool designed to automate the creation of training datasets for software architecture traceability research. It leverages automated web-mining of technical libraries and big-data analysis of over 22 million source files to identify and extract implementation samples of architectural tactics.
TL;DR
The manual collection of training data is the "Achilles' heel" of automated software traceability. BUDGET (Big-data aUgmented Dataset GEnerator Tool) solves this by mining technical libraries and ultra-large-scale code repositories (22M+ files) to automatically generate high-quality training sets for architectural tactics with over 90% accuracy.
The Bottleneck: Why Traceability Research is Stalled
Software architecture traceability—the ability to link requirements to specific architectural tactics and code—is vital for safety-critical systems and compliance (e.g., HIPAA). However, supervised learning models for this task require massive amounts of labeled data.
Currently, a researcher might spend three months manually vetting code snippets just to train a classifier for five or ten tactics. This "manual bottleneck" prevents the community from scaling research to complex industrial systems and limits the diversity of programming languages and domains studied.
Methodology: Dual-Engine Data Harvesting
BUDGET takes a two-pronged approach to bypass the manual labor of dataset creation:
1. Web-Mining Technical Knowledge
The tool doesn't just look for keywords; it uses textbook definitions of architectural tactics (like Heartbeat or Resource Pooling) to construct intelligent queries via the Google Search API. It targets high-authority technical sites like MSDN and Oracle Documentation to extract "clean" specifications and API usage examples.
2. Big-Data Code Analysis
For implementation samples, BUDGET mines an astronomical repository of 116,609 projects from GitHub, SourceForge, and Apache.
- Indexing: It uses a custom-built infrastructure that stems and fingerprints 22 million files.
- Search: It employs a parallelized Vector Space Model (VSM) to calculate cosine similarity between the tactic's textual description and the implementation code.

Performance & Accuracy
The authors validated the tool's output against human peer review. The results demonstrate that automated mining is not just faster, but highly reliable:
- Accuracy: The Big-Data engine achieved 100% accuracy for tactics like Audit, Scheduling, and Authentication.
- Speed: Searching across the 22-million-file index takes only a few seconds.
- Diversity: By mining GitHub and Apache, the tool avoids the "silo effect" of single-project datasets.

Critical Analysis & Conclusion
BUDGET represents a transition toward Evidence-Based Software Engineering. Its ability to generate negative samples—code that looks like a tactic but isn't—is particularly crucial for training robust classifiers that don't produce excessive false positives.
Limitations:
- While the tool supports 10 hard-coded tactics, user-defined queries require careful keyword selection to maintain high precision.
- The Web-Mining approach for "Heartbeat" and "Audit" showed lower accuracy (60%), likely due to the commonality of these terms in generic IT documentation.
Looking Ahead: BUDGET paves the way for "Self-Evolving Research." As GitHub grows, BUDGET's knowledge base grows with it. Future iterations could potentially integrate Deep Learning (e.g., CodeBERT) to further refine implementation detection beyond simple VSM similarity.
