Harmony: Breaking the Interference Barrier in ML Clusters with Deep Reinforcement Learning

Deep Learning-based Job Placement in Distributed Machine Learning Clusters

2019-04-01
Yixin Bao, Yanghua Peng, Chuan Wu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Harmony, a Deep Reinforcement Learning (DRL)-driven scheduler for distributed Machine Learning (ML) clusters. It optimizes job placement by accounting for performance interference among co-located workloads, achieving a 25% reduction in average job completion time compared to state-of-the-art schedulers like Kubernetes and Tetris.

TL;DR

In modern distributed Machine Learning (ML) clusters, co-locating jobs is essential for utilization but often leads to "invisible" performance interference. Harmony is a DRL-based scheduler that learns to place jobs where they interfere least. By doubling down on an auxiliary reward model to simulate training data, it achieves a 25% speedup in job completion times over standard industry schedulers like Kubernetes.

The "Invisible" Cost of Co-location

As production ML clusters scale, the go-to strategy for efficiency is multi-resource bin packing—cramming as many jobs as possible onto a server based on CPU and memory limits.

However, the authors reveal a stark reality: Case Study 1 showed that co-locating standard jobs (like ResNet and VGG) can cause a nearly 2x slowdown for specific models. Why? Because existing schedulers are "interference-oblivious." They don't see the silent battles for the PCIe bus, CPU caches, or disk I/O.

Traditional "white-box" solutions attempt to build mathematical models for this. But distributed ML is too dynamic—a CPU-intensive CTC model affects a network-heavy VGG-16 model differently than it would a memory-bound Seq2Seq model.

Methodology: The Harmony Architecture

Harmony moves away from manual profiling and instead treats job placement as a Reinforcement Learning problem.

1. The DRL Agent (The Brain)

The system uses an Actor-Critic framework to map the "raw" state of the cluster (available resources, demands of new jobs, and current placement of active jobs) to an action: Which server should host this specific Worker or Parameter Server (PS)?

2. Solving the "Cold Start" and Data Scarcity Problem

Standard RL requires millions of trials to converge. You cannot afford to crash a production cluster for weeks just to "train" a scheduler. Harmony’s Innovation: The authors built an Auxiliary Reward Prediction Model.

  • First, they take a small set of historical traces.
  • They train a supervised Neural Network to predict "Training Speed" given a specific placement.
  • This model acts as a "simulator," allowing the DRL agent to "hallucinate" millions of placement scenarios and learn from them without touching a single real server.

3. Stabilizing the Policy

To ensure the DRL doesn't get stuck in local optima, the authors implemented:

  • Experience Replay: Breaking the correlation between consecutive scheduling intervals.
  • Job-Aware Exploration: Mixing standard RL exploration with "expert" heuristics like Load Balancing and Bin Packing (-greedy) to guide the agent toward sensible baselines early on.

Harmony DRL Architecture

Experiments & Results: Real-World Gains

The authors didn't just simulate; they deployed Harmony on a real Kubernetes GPU cluster with 1080Ti GPUs and MXNet workloads.

  • Performance Boost: Harmony reduced average Job Completion Time (JCT) by 25% compared to Kubernetes' default Load Balancing and the resource-optimizing Tetris (multi-resource bin packing).
  • Accuracy: Their Reward Prediction Model achieved a 9.8% relative error, significantly outperforming traditional linear interference models (which failed to account for non-CPU resources like GPU bus contention).

Performance Comparison

Critical Analysis & Conclusion

Takeaway

Harmony proves that the "Black-box" approach—learning interference implicitly through a Neural Network—is more robust and extensible than "White-box" analytical models. The transition from manually crafted heuristics to learned policies is the next frontier for cluster management.

Limitations & Future Work

One current limitation is that Harmony assumes job placement is static once a job starts. In highly dynamic tiers, live migration or dynamic scaling of workers could further optimize performance. Additionally, as clusters move toward specialized architectures (NVLink, RDMA), the state space will grow, requiring even more sophisticated embedding of the cluster topology into the DRL's state representation.

The bottom line: In the race for AI efficiency, your choice of scheduler is just as important as your choice of hardware.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Deep Reinforcement Learning for GPU-specific resource scheduling in multi-tenant clusters, specifically focusing on fragmentation and interconnect contention.
  • Identify the origin of the "proxy reward model" or "world model" concept in reinforcement learning and how subsequent systems like Harmony improved its accuracy for systems tasks.
  • Explore how DRL-based scheduling methods have been extended to large-scale, heterogeneous clusters involving both CPUs, GPUs, and specialized AI accelerators like TPUs or NPUs.
Contents
Harmony: Breaking the Interference Barrier in ML Clusters with Deep Reinforcement Learning
1. TL;DR
2. The "Invisible" Cost of Co-location
3. Methodology: The Harmony Architecture
3.1. 1. The DRL Agent (The Brain)
3.2. 2. Solving the "Cold Start" and Data Scarcity Problem
3.3. 3. Stabilizing the Policy
4. Experiments & Results: Real-World Gains
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work