Beyond User Guesses: Leveraging Job Metadata for Precision HPC Scheduling

Analysis of Job Metadata for Enhanced Wall Time Prediction

2019-01-01
Mehmet Soysal, Marco Berghoff, Achim Streit
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an automated machine learning approach to improve job wall time prediction in High-Performance Computing (HPC) systems. Using the auto-sklearn library, the authors demonstrate that incorporating previously overlooked metadata, such as job names and initial working directories (IWD), significantly outperforms traditional user-provided estimates.

TL;DR

Predicting how long a supercomputing job will run is notoriously difficult because users tend to "pad" their estimates. This paper introduces an AutoML-based framework that analyzes overlooked metadata—specifically working directories and job names—to predict wall times. The result? A staggering improvement where prediction errors drop from over 7 hours to just about 1 hour.

The "User Padding" Problem

In HPC environments, the scheduler is the conductor of a massive orchestra. To do its job, it needs to know how long each "piece" (job) will play. Currently, schedulers rely on user-provided wall time estimates. However, users face a perverse incentive: if they underestimate, the system kills their job. Consequently, they overestimate wildly.

Prior research attempted to fix this using templates or basic job features like core counts. But as the researchers at KIT point out, these methods miss the contextual fingerprints of a job—where it is running and what it is called.

Methodology: Extracting Signal from Strings

The authors argue that the "Initial Working Directory" (IWD) and "Jobname" are not just strings; they are indicators of intent. A job running in /sim/run_a is likely different from one in /test/debug.

1. Feature Decomposition

The team applied regular expressions to break down these strings into matrices.

  • IWD Analysis: Separated the file system type (e.g., home vs. scratch) from the project and sub-directories.
  • Temporal Logic: They extracted the "Hour of Day" and "Day of Week" from submission timestamps, identifying patterns where "weekend jobs" or "overnight runs" behave differently.

2. The AutoML Pipeline

Instead of manually tuning a single regressor, they used auto-sklearn. This allows the system to automatically explore various pre-processors (like PCA) and models (like Random Forests or Gradient Boosting) to find the best fit for each individual user's specific habits.

Model Feature Categorization Figure 1: Comparison of two users. User A's jobs are predictable via task count; User B requires more complex metadata to distinguish runtimes.

Experimental Results: A 7x Accuracy Boost

The study evaluated traces from the ForHLR I and II clusters, involving hundreds of users and over 270,000 jobs.

  • The R² Metric: While 80-90% of users provide estimates so poor they lead to negative scores, the AutoML models shifted the distribution toward excellence (scores > 0.8).
  • Error Reduction: The Mean Absolute Error (MAE) for ForHLR II dropped from 16 hours (user estimate) to 3 hours (AutoML).

Accuracy Distribution (R² Score) Figure 2: Cumulative distribution of scores. The "ALL" line (rightmost) shows that combining all metadata yields the highest prediction density near a perfect score of 1.0.

Deep Insight: Why Directories Matter

The "Curse of Dimensionality" is a threat here—if a user creates a new directory for every single job, the data becomes too sparse to learn. However, the study found that most users are creatures of habit. They use structured naming conventions that serve as an implicit "tag" for the computational complexity of the task. By mining these tags, the model effectively performs automated profiling without ever looking at the actual source code or input data.

Conclusion & Future Impact

This research moves us closer to "Transparent Scheduling," where the system understands the workload better than the user does.

  • Limitations: The "Cold Start" problem remains; the system needs historical data for a user before it can predict their behavior accurately.
  • Outlook: Integrating these models directly into schedulers like Slurm or PBS could allow for "Planning Wall Times" that coexist with "Guaranteed Wall Times," drastically increasing the utilization and throughput of next-generation Exascale systems.

Author's Perspective: This work proves that in technical systems, the context (metadata) is often as valuable as the content (payload). For HPC administrators, the message is clear: stop asking users for better estimates and start mining the data you already have.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Deep Learning or Transformer-based architectures to predict HPC job wall times using the Standard Workload Format (SWF).
  • Which study first identified the "user overestimation bias" in HPC scheduling, and how do modern dynamic backfilling algorithms mitigate this today?
  • Research how feature engineering techniques for path/directory strings in HPC metadata can be improved using natural language processing (NLP) embeddings instead of simple regular expressions.
Contents
Beyond User Guesses: Leveraging Job Metadata for Precision HPC Scheduling
1. TL;DR
2. The "User Padding" Problem
3. Methodology: Extracting Signal from Strings
3.1. 1. Feature Decomposition
3.2. 2. The AutoML Pipeline
4. Experimental Results: A 7x Accuracy Boost
5. Deep Insight: Why Directories Matter
6. Conclusion & Future Impact