Beyond User Guesses: Leveraging Job Metadata for Precision HPC Scheduling
Analysis of Job Metadata for Enhanced Wall Time Prediction
This paper presents an automated machine learning approach to improve job wall time prediction in High-Performance Computing (HPC) systems. Using the auto-sklearn library, the authors demonstrate that incorporating previously overlooked metadata, such as job names and initial working directories (IWD), significantly outperforms traditional user-provided estimates.
TL;DR
Predicting how long a supercomputing job will run is notoriously difficult because users tend to "pad" their estimates. This paper introduces an AutoML-based framework that analyzes overlooked metadata—specifically working directories and job names—to predict wall times. The result? A staggering improvement where prediction errors drop from over 7 hours to just about 1 hour.
The "User Padding" Problem
In HPC environments, the scheduler is the conductor of a massive orchestra. To do its job, it needs to know how long each "piece" (job) will play. Currently, schedulers rely on user-provided wall time estimates. However, users face a perverse incentive: if they underestimate, the system kills their job. Consequently, they overestimate wildly.
Prior research attempted to fix this using templates or basic job features like core counts. But as the researchers at KIT point out, these methods miss the contextual fingerprints of a job—where it is running and what it is called.
Methodology: Extracting Signal from Strings
The authors argue that the "Initial Working Directory" (IWD) and "Jobname" are not just strings; they are indicators of intent. A job running in /sim/run_a is likely different from one in /test/debug.
1. Feature Decomposition
The team applied regular expressions to break down these strings into matrices.
- IWD Analysis: Separated the file system type (e.g., home vs. scratch) from the project and sub-directories.
- Temporal Logic: They extracted the "Hour of Day" and "Day of Week" from submission timestamps, identifying patterns where "weekend jobs" or "overnight runs" behave differently.
2. The AutoML Pipeline
Instead of manually tuning a single regressor, they used auto-sklearn. This allows the system to automatically explore various pre-processors (like PCA) and models (like Random Forests or Gradient Boosting) to find the best fit for each individual user's specific habits.
Figure 1: Comparison of two users. User A's jobs are predictable via task count; User B requires more complex metadata to distinguish runtimes.
Experimental Results: A 7x Accuracy Boost
The study evaluated traces from the ForHLR I and II clusters, involving hundreds of users and over 270,000 jobs.
- The R² Metric: While 80-90% of users provide estimates so poor they lead to negative scores, the AutoML models shifted the distribution toward excellence (scores > 0.8).
- Error Reduction: The Mean Absolute Error (MAE) for ForHLR II dropped from 16 hours (user estimate) to 3 hours (AutoML).
Figure 2: Cumulative distribution of scores. The "ALL" line (rightmost) shows that combining all metadata yields the highest prediction density near a perfect score of 1.0.
Deep Insight: Why Directories Matter
The "Curse of Dimensionality" is a threat here—if a user creates a new directory for every single job, the data becomes too sparse to learn. However, the study found that most users are creatures of habit. They use structured naming conventions that serve as an implicit "tag" for the computational complexity of the task. By mining these tags, the model effectively performs automated profiling without ever looking at the actual source code or input data.
Conclusion & Future Impact
This research moves us closer to "Transparent Scheduling," where the system understands the workload better than the user does.
- Limitations: The "Cold Start" problem remains; the system needs historical data for a user before it can predict their behavior accurately.
- Outlook: Integrating these models directly into schedulers like Slurm or PBS could allow for "Planning Wall Times" that coexist with "Guaranteed Wall Times," drastically increasing the utilization and throughput of next-generation Exascale systems.
Author's Perspective: This work proves that in technical systems, the context (metadata) is often as valuable as the content (payload). For HPC administrators, the message is clear: stop asking users for better estimates and start mining the data you already have.
