Scheduling multithreaded computations by work stealing

doi:10.1145/324133.324234

Journal ArticleDOI

Scheduling multithreaded computations by work stealing

Robert D. Blumofe, +1 more

- 01 Sep 1999 -

Journal of the ACM

- Vol. 46, Iss: 5, pp 720-748

TLDR

This paper gives the first provably good work-stealing scheduler for multithreaded computations with dependencies, and shows that the expected time to execute a fully strict computation on P processors using this scheduler is 1:1.

Abstract:

This paper studies the problem of efficiently schedulling fully strict (i.e., well-structured) multithreaded computations on parallel computers. A popular and practical method of scheduling this kind of dynamic MIMD-style computation is “work stealing,” in which processors needing work steal computational threads from other processors. In this paper, we give the first provably good work-stealing scheduler for multithreaded computations with dependencies.Specifically, our analysis shows that the expected time to execute a fully strict computation on P processors using our work-stealing scheduler is T1/P + O(T ∞ , where T1 is the minimum serial execution time of the multithreaded computation and (T ∞ is the minimum execution time with an infinite number of processors. Moreover, the space required by the execution is at most S1P, where S1 is the minimum serial space requirement. We also show that the expected total communication of the algorithm is at most O(PT ∞( 1 + nd)Smax), where Smax is the size of the largest activation record of any thread and nd is the maximum number of times that any thread synchronizes with its parent. This communication bound justifies the folk wisdom that work-stealing schedulers are more communication efficient than their work-sharing counterparts. All three of these bounds are existentially optimal to within a constant factor.

Scheduling multithreaded computations by work stealing

Citations

Language run-time systems: An overview

Skueue: A Scalable and Sequentially Consistent Distributed Queue

Automatic task and data mapping in shared memory architectures

Performance Analysis of Work Stealing for Streaming Systems and Optimizations

Toward Better Computation Models for Modern Machines

References

Bounds on Multiprocessing Timing Anomalies

Cilk: An Efficient Multithreaded Runtime System

Bounds for certain multiprocessing anomalies

The implementation of the Cilk-5 multithreaded language

The Parallel Evaluation of General Arithmetic Expressions

Related Papers (5)

The implementation of the Cilk-5 multithreaded language

X10: an object-oriented approach to non-uniform cluster computing

Cilk: an efficient multithreaded runtime system

A Java fork/join framework

MapReduce: simplified data processing on large clusters