MGPUSim: enabling multi-GPU performance modeling and optimization

doi:10.1145/3307650.3322230

Open AccessProceedings ArticleDOI

MGPUSim: enabling multi-GPU performance modeling and optimization

Yifan Sun, +17 more

- pp 197-209

Chats0

TLDR

This work presents MGPUSim, a cycle-accurate, extensively validated, multi-GPU simulator, based on AMD's Graphics Core Next 3 (GCN3) instruction set architecture, and proposes the Locality API, an API extension that allows the GPU programmer to both avoid the complexity of multi- GPU programming, while precisely controlling data placement in the multi- GPUs memory.

Abstract:

The rapidly growing popularity and scale of data-parallel workloads demand a corresponding increase in raw computational power of Graphics Processing Units (GPUs). As single-GPU platforms struggle to satisfy these performance demands, multi-GPU platforms have started to dominate the high-performance computing world. The advent of such systems raises a number of design challenges, including the GPU microarchitecture, multi-GPU interconnect fabric, runtime libraries, and associated programming models. The research community currently lacks a publicly available and comprehensive multi-GPU simulation framework to evaluate next- generation multi-GPU system designs. In this work, we present MGPUSim, a cycle-accurate, extensively validated, multi-GPU simulator, based on AMD's Graphics Core Next 3 (GCN3) instruction set architecture. MGPUSim comes with in-built support for multi-threaded execution to enable fast, parallelized, and accurate simulation. In terms of performance accuracy, MGPUSim differs by only 5.5% on average from the actual GPU hardware. We also achieve a 3.5x and a 2.5x average speedup running functional emulation and detailed timing simulation, respectively, on a 4-core CPU, while delivering the same accuracy as serial simulation. We illustrate the flexibility and capability of the simulator through two concrete design studies. In the first, we propose the Locality API, an API extension that allows the GPU programmer to both avoid the complexity of multi-GPU programming, while precisely controlling data placement in the multi-GPU memory. In the second design study, we propose Progressive Page Splitting Migration (PASI), a customized multi-GPU memory management system enabling the hardware to progressively improve data placement. For a discrete 4-GPU system, we observe that the Locality API can speed up the system by 1.6x (geometric mean), and PASI can improve the system performance by 2.6x (geometric mean) across all benchmarks, compared to a unified 4-GPU platform.

MGPUSim: enabling multi-GPU performance modeling and optimization

Citations

Accel-sim: an extensible simulation framework for validated GPU modeling

Characterizing and Modeling Non-Volatile Memory Systems

Towards Inspecting and Eliminating Trojan Backdoors in Deep Neural Networks

AccelWattch: A Power Modeling Framework for Modern GPUs

Grus: Toward Unified-memory-efficient High-performance Graph Processing on GPU

References

ImageNet Classification with Deep Convolutional Neural Networks

ImageNet: A large-scale hierarchical image database

ImageNet classification with deep convolutional neural networks

Parallel discrete event simulation

Analyzing CUDA workloads using a detailed GPU simulator

Related Papers (5)

Analyzing CUDA workloads using a detailed GPU simulator

Hetero-mark, a benchmark suite for CPU-GPU collaborative computing

Multi2Sim: a simulation framework for CPU-GPU computing

MCM-GPU: Multi-Chip-Module GPUs for Continued Performance Scalability

Parallel GPU architecture simulation framework exploiting work allocation unit parallelism