MSR4ML: Reconstructing the Missing Links in Machine Learning Repositories
MSR4ML: Reconstructing Artifact Traceability in Machine Learning Repositories
The paper introduces MSR4ML, an automated framework that leverages Mining Software Repositories (MSR) and static code analysis to reconstruct traceability links between code, data, and models in ML projects. It aims to bridge the gap in Git-based versioning by automatically mapping artifact evolution through commit history.
TL;DR
Integrating Machine Learning into traditional software engineering workflows often breaks artifact traceability. MSR4ML is a new framework that uses static code analysis and AST parsing to automatically find the hidden links between your Python code, your data files, and your trained models. It achieves over 84% precision in identifying these relationships without requiring developers to manually tag their files.
Background: The ML Traceability Gap
In standard software, Git is king. But in ML, the "source of truth" is split across code, massive datasets, and binary model weights. Standard Git handles code well but ignores the rest. When a model's performance drops, developers are often left asking: "Which commit changed the data preprocessing logic that affected this specific .pkl model?"
Current solutions like DVC (Data Version Control) solve this but require significant manual setup. MSR4ML targets the "late development stage" where a project is already complex, and manual tagging is no longer feasible.
Methodology: How MSR4ML Sees Your Project
The framework operates as a pipeline of four specialized modules designed to turn a flat Git repository into a weighted graph of dependencies.
1. The Core Architecture
The system doesn't just look for filenames; it understands intent by parsing the Abstract Syntax Tree (AST).

2. Deepening the Hunt: Value Inference
A major challenge in static analysis is resolving variables. If a script calls pd.read_csv(data_path), the tool must figure out what data_path actually points to. MSR4ML uses Argument Value Inference to trace variable assignments back through the AST, resolving string literals and concatenations to find the real file on disk.
3. Classification and Weights
Not all links are equal. The framework assigns weights based on proximity:
- Direct Access: A script that reads a CSV has a high-weight link.
- Coupled Access: A script that calls a function in another script which then reads the CSV has a lower-weight link. This allows the "Commit Tracker" to prioritize which code changes are most likely responsible for a model's current state.
Experimental Results: Precision vs. Recall
The researchers tested MSR4ML against 20 repositories from the "Papers With Code" dataset.

- Precision (84.1%): The tool is highly reliable at identifying correct artifacts, avoiding "spamming" the developer with false positives.
- Recall (61.3%): While lower, the authors noted that most missed artifacts were due to complex path logic that static analysis couldn't resolve—a gap they plan to fill with dynamic analysis.
- Impact of Inference: Adding specialized inference functions for the
oslibrary alone boosted recall by 35%, highlighting that the secret sauce is in how well the tool understands Python's standard library.
Critical Insight: Why This Matters
The real value of MSR4ML isn't just knowing what is in a repo, but enabling Model-Aware Git Blame.
Imagine running a command that doesn't just show you who changed a line of code, but who changed the hyperparameters or the training set that led to a specific model version. By reconstructing these links automatically, MSR4ML moves us closer to a "Self-Organizing MLOps" environment where the infrastructure understands the model lifecycle as deeply as the developer does.
Limitations & Future Work
The current prototype is limited to static analysis, meaning it struggles with dynamically generated file paths (e.g., paths loaded from a config file at runtime). The authors propose adding Dynamic Code Analysis in the next iteration to execute code blocks and capture the actual file handles in real-time.
Conclusion
MSR4ML provides a compelling roadmap for turning "Black Box" ML repositories into transparent, traceable software systems. For teams struggling with technical debt in their ML pipelines, this represents a significant step toward automated governance and reproducibility.
