Predicting High-Performance Computing Job Performance: A Machine Learning Approach
Machine Learning Based Performance Analysis and Prediction of Jobs on a HPC Cluster
This paper presents a machine learning-based framework for predicting the performance (CPU time) of parallel jobs on an HPC cluster. By leveraging historical job logs from Northwestern Polytechnical University, the authors compare multiple models—Multivariate Linear Regression, Polynomial Regression, and Neural Networks—achieving a prediction accuracy of over 83% for specific user applications.
TL;DR
Researchers at Northwestern Polytechnical University have developed a framework to predict the CPU time of HPC jobs by mining years of historical logs. By training specialized models (Linear, Polynomial, and Neural Networks) per user and application, they successfully predicted the performance of new scientific jobs (like VASP) with over 83% accuracy, even when critical memory features were initially unknown.
Background & Positioning
In the world of High-Performance Computing (HPC), job scheduling is the "brain" that determines efficiency. Most schedulers rely on user-estimated walltimes, which are often exaggerated to prevent jobs from being killed. This work moves away from user intuition and toward Data-Driven Performance Modeling, situating itself as a practical application of regression and neural networks to real-world infrastructure management.
The Problem: The "Informational Gap"
Existing methods fail primarily because:
- Data Complexity: The relationship between requested cores, memory, and actual CPU time is rarely linear.
- Missing Features: Some of the most predictive features, such as
UsedMemory, are only known after a job completes, making them useless for pre-execution prediction unless they can be estimated.
The authors' insight was to stop looking for a universal model and instead focus on user-specific behaviors and application-specific "seeds" (found in input files) to predict these missing features.
Methodology: Specialized Model Architecture
The workflow involves rigorous data cleaning of Torque-based logs, followed by the deployment of four distinct model types:
- Multivariate Linear Regression: Establishing a baseline by assuming linear weights for features like
StartTimeandCoreNumber. - Multivariate Polynomial Regression: Capturing non-linearities using Taylor expansion-style terms (up to degree 4).
- Linear Neural Networks: A single-layer approach for fast iteration.
- BP Neural Networks (Back-Propagation): A three-layer (input, hidden, output) architecture designed to find complex, hidden correlations.
Bridging the Feature Gap
To solve the "missing feature" problem for new VASP jobs, the authors extracted the Mesh Matrix from VASP input files. They discovered that the computational volume () calculated from these matrices significantly correlates with memory usage.
Table 1: The 10 key features extracted from job logs used for training.
Experiments & Core Insights
The researchers found an "Overfitting Threshold" in polynomial models. While increasing the degree reduced training error, it often caused spikes in testing error (e.g., degree 4 for user xuepy led to an MSE of 39.979 compared to 1.576 at degree 2).
Performance Comparison
The results confirm the necessity of "Model Selection per User":
- User USPEX: BP Neural Network was the clear winner (MSE 3.638).
- User xuepy: Polynomial Regression (Degree 2) was most effective (MSE 1.576).
Figure 1: Distribution of measured (red) vs. predicted (blue) values for specific users.
Critical Analysis & Conclusion
Takeaway
The study proves that even with "messy" real-world logs, ML can achieve high accuracy if we leverage application-specific domain knowledge (like parsing VASP input files). This shifts HPC management from reactive to proactive.
Limitations
- Feature Sensitivity: The dependency on specific input file formats (like VASP's OUTCAR) makes the method difficult to generalize to "black-box" proprietary software.
- Static Nature: The models are trained on historical data and may need frequent retraining as cluster hardware or compiler versions change.
Future Work
The authors aim to extend this to GPU-based jobs, where the performance bottlenecks (PCIe bandwidth, VRAM) are significantly different from CPU-bound tasks. This research paves the way for "Intelligent Schedulers" that can automatically adjust job priority based on predicted resource footprints.
