Merge or Separate? Optimizing Heterogeneous Scheduling via Machine Learning

Merge or Separate?: Multi-job Scheduling for OpenCL Kernels on CPU/GPU Platforms

2017-02-04
Yuan Wen, Michael F.P. O'Boyle, M. O’Boyle
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a runtime framework and JIT compiler for multi-job scheduling of OpenCL kernels on heterogeneous CPU/GPU platforms. It utilizes machine learning-based predictive models to decide whether to merge kernels for concurrent GPU execution or separate them across CPU and GPU devices, achieving SOTA performance in multi-user environments.

TL;DR

As heterogeneous computing becomes mainstream, the "one GPU per application" model is breaking down. This paper presents a runtime framework that uses machine learning to dynamically decide whether to run OpenCL kernels concurrently on a GPU or offload them to the CPU. By combining a JIT compiler with predictive modeling, it improves throughput by up to 36% without requiring any prior profiling of applications.

The Resource Underutilization Trap

While GPUs are powerhouses for parallel computation, they are often surprisingly idle. Many kernels don't fully saturate the GPU's memory bandwidth or execution units. In a multi-user environment, simply running tasks "First-Come-First-Served" (FCFS) leads to two massive inefficiencies:

  1. GPU Underutilization: Small kernels leave much of the hardware dormant.
  2. Wasted Host Power: The multi-core CPU, capable of handling specific workloads more efficiently than a GPU (especially those with high data transfer overhead), often sits idle.

The industry has tried to solve this via "Elastic Kernels" or "Energy-Efficient Concurrent Kernels" (EECK), but these suffer from a "profiling tax"—you have to run the code beforehand to know how it behaves. In a dynamic cloud or multi-tasking OS environment, you don't have that luxury.

Methodology: The "At the Factory" Intelligence

The core insight of this research is that while individual characteristics (like branch divergence or memory intensity) are poor predictors of performance on their own, a statistical combination of these features is highly accurate.

1. Feature Extraction & JIT Merging

The framework acts as a transparent layer. When a user submits an OpenCL kernel, the JIT compiler extracts:

  • Static Features: Instruction counts (math, memory, branches, barriers).
  • Dynamic Features: Global/local work sizes and data transfer sizes.

For kernels that the model decides to "merge," the JIT compiler uses an improved version of inter-thread-block-fusion. It employs a modulo operation to partition Computing Units (CUs) and thread decoupling to ensure the GPU treats the merged kernels as a single, resource-efficient entity.

Overall Framework Architecture

2. The Predictive Brain

The authors use two levels of Decision Trees:

  • Model A: Determines device affinity (Should this go to the CPU or GPU?).
  • Model B: Determines merging suitability (If it's for the GPU, should it pair with another kernel? At what resource ratio?).

Experiments: Breaking the SOTA

The system was tested on both NVIDIA (GTX 780) and AMD (Radeon HD7970) platforms using Parboil and Polybench benchmarks.

Performance Gains

The approach, dubbed SoC (Separate or Concurrent), was compared against "Heterogeneous Scheduling" (HS) and the profiling-heavy EECK.

  • Throughput: SoC outperformed HS by 17% on NVIDIA and 10% on AMD.
  • Turnaround Time: The improvements in ANTT (Average Normalized Turnaround Time) were even more dramatic, reaching up to 52% improvement on NVIDIA. This suggests the model is excellent at identifying kernels that would otherwise "choke" the GPU and rerouting them to the CPU.

Performance results on NVIDIA and AMD

Deep Insight: Why Simple Heuristics Fail

The paper provides a fascinating ablation study on kernel characteristics. For instance, the common wisdom that "co-running high-compute and high-memory intensity kernels is best" is debunked. The authors show that co-running two low-compute kernels can actually yield a 7% gain, while co-running high-compute with low-compute often causes a 35% slowdown due to resource contention. This justifies the necessity of the machine learning approach over manual tuning.

Critical Analysis & Future Outlook

While the results are impressive, the framework has a few limitations:

  • Binary Merging: It currently only merges up to two kernels. In systems with massive resource counts (like NVIDIA H100s), merging 3 or 4 kernels might be necessary.
  • Static Scaling: The model is trained "at the factory." While it is portable, its accuracy might dip on radically new architectures (e.g., Apple Silicone's Unified Memory or Intel's XPU).

Conclusion: This work proves that the "Merge or Separate" decision is the new frontier of GPGPU scheduling. By leveraging a low-overhead JIT compiler and a smart predictive model, we can treat the CPU and GPU as a cohesive pool of resources rather than isolated islands.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize deep learning or reinforcement learning instead of decision trees for dynamic OpenCL kernel scheduling on heterogeneous systems.
  • Which paper first introduced the concept of inter-thread-block-fusion for GPU kernel merging, and how does the current work's implementation differ in handling workgroup indices?
  • Find studies that apply these machine learning-based scheduling techniques to modern unified memory architectures or Multi-Instance GPU (MIG) environments.
Contents
Merge or Separate? Optimizing Heterogeneous Scheduling via Machine Learning
1. TL;DR
2. The Resource Underutilization Trap
3. Methodology: The "At the Factory" Intelligence
3.1. 1. Feature Extraction & JIT Merging
3.2. 2. The Predictive Brain
4. Experiments: Breaking the SOTA
4.1. Performance Gains
5. Deep Insight: Why Simple Heuristics Fail
6. Critical Analysis & Future Outlook