PISCES: Breaking the Synchronization Barrier in Multi-Job MapReduce

PISCES: Optimizing Multi-Job Application Execution in MapReduce

2016-08-26
Qi Chen, Jinyu Yao, Benchao Li, Zhen Xiao
Summary
Problem
Method
Results
Takeaways
Abstract

PISCES is a specialized optimization framework for multi-job MapReduce applications that introduces inter-job data pipelining and critical chain scheduling. It achieves SOTA performance by breaking the synchronization barriers between dependent jobs, improving execution speed by up to 52%.

TL;DR

PISCES (Pipeline Improvement Support with Critical chain Estimation Scheduling) is an architectural extension for MapReduce designed to accelerate multi-job workflows. It eliminates the "wait-until-finished" bottleneck between dependent jobs through an innovative data-pipelining mechanism and a sophisticated critical-chain scheduler driven by Locally Weighted Linear Regression (LWLR).

The "Synchronization Barrier" Problem

In modern big data stacks, a single high-level query (e.g., a Hive or Pig script) is often decomposed into a Directed Acyclic Graph (DAG) of multiple MapReduce jobs. Standard Hadoop, however, is dependency-blind.

Current systems suffer from two major inefficiencies:

  1. Blockage: Job B cannot start its Map phase until Job A has completely finished and moved its output to a final HDFS directory.
  2. Ignorance: Schedulers view jobs as a flat list, often failing to prioritize the "Critical Chain"—the sequence of jobs that actually dictates the total completion time.

Methodology: The PISCES Architecture

PISCES introduces three key modules into the MapReduce Application Master: a Dependency Analyzer, a Job Time Estimator, and a Job Scheduler.

1. Inter-Job Data Pipelining

The most radical shift in PISCES is the ability to stream data between jobs. By tapping into the temporary output directories of HDFS, PISCES allows downstream Map tasks to spawn as soon as an upstream Reduce task flushes a single data block (64MB).

To ensure correctness and fault tolerance, PISCES utilizes a Hard Link feature in HDFS. This ensures that even if an upstream job finishes and tries to move its temporary files, the downstream Map tasks still have a valid reference to the data blocks.

PISCES Pipeline Process

2. Critical Chain Scheduling with LWLR

Not all jobs are created equal. PISCES uses Locally Weighted Linear Regression (LWLR) to predict job runtimes based on historical data. Unlike traditional linear models, LWLR gives more weight to recent and "size-similar" job executions, allowing it to accurately predict both linear and super-linear (e.g., ) workloads.

The scheduler then identifies the Critical Chain—the longest path of execution in the DAG—and ensures these jobs receive priority resource allocation to minimize the total Makespan.

Performance Benchmarks

The authors tested PISCES using PageRank (iterative) and PigMix (database queries) workloads.

  • Parallelism Boost: PISCES increased the degree of system parallelism by 68% in database operations.
  • Speedup: Iterative applications like PageRank ran 41% faster because the system effectively eliminated the "dead time" between iterations.
  • Resource Efficiency: By pipelining data, PISCES kept more data in the OS file cache, reducing expensive disk I/O and swap operations compared to standard Pig.

Performance Comparison in PigMix

Critical Insight & Conclusion

The genius of PISCES lies in its "bottom-up" approach. While tools like Pig and Hive try to optimize the logical plan, PISCES optimizes the physical execution by making the MapReduce engine itself dependency-aware.

Takeaway: For any system managing job graphs, the ability to overlap the "Producer's Finish" with the "Consumer's Start" is the single most effective way to reclaim lost cluster cycles. While PISCES was built for MapReduce, its logic of LWLR-based estimation and hard-link-enabled pipelining remains highly relevant for modern cloud-native orchestrators.

Limitations

  • Cold Start: The LWLR estimator requires historical data to be accurate; performance on unique, one-off jobs may revert to standard heuristics.
  • HDFS Modifications: The requirement for "Hard Link" support in HDFS means PISCES isn't a "plug-and-play" library but requires a modified filesystem layer.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend inter-job pipelining concepts to modern Spark or Flink architectures for iterative graph processing.
  • Which study first introduced the concept of Critical Path Method (CPM) in distributed task scheduling, and how does PISCES's Critical Chain differ from it?
  • Find research evaluating the impact of Locally Weighted Linear Regression (LWLR) versus Deep Learning models for predicting execution times in heterogeneous cloud environments.
Contents
PISCES: Breaking the Synchronization Barrier in Multi-Job MapReduce
1. TL;DR
2. The "Synchronization Barrier" Problem
3. Methodology: The PISCES Architecture
3.1. 1. Inter-Job Data Pipelining
3.2. 2. Critical Chain Scheduling with LWLR
4. Performance Benchmarks
5. Critical Insight & Conclusion
5.1. Limitations