(PDF) A scalable processing-in-memory accelerator for parallel graph processing (2015) | Junwhan Ahn

Citations

PDF

Open Access

More filters

Proceedings Article•DOI•

Toward standardized near-data processing with unrestricted data placement for GPUs

[...]

Gwangsun Kim, Niladrish Chatterjee¹, Mike O'Connor², Kevin Hsieh³•Institutions (3)

Nvidia¹, University of Texas at Austin², Carnegie Mellon University³

12 Nov 2017

TL;DR: By enhancing this system with a smart offload selection mechanism that is cognizant of the compute capability of the NDP and cache locality on the host processor, system performance and energy are improved by up to 66.8% and 37.6%, respectively.

...read moreread less

Abstract: 3D-stacked memory devices with processing logic can help alleviate the memory bandwidth bottleneck in GPUs. However, in order for such Near-Data Processing (NDP) memory stacks to be used for different GPU architectures, it is desirable to standardize the NDP architecture. Our proposal enables this standardization by allowing data to be spread across multiple memory stacks as is the norm in high-performance systems without an MMU on the NDP stack. The keys to this architecture are the ability to move data between memory stacks as required for computation, and a partitioned execution mechanism that offloads memory-intensive application segments onto the NDP stack and decouples address translation from DRAM accesses. By enhancing this system with a smart offload selection mechanism that is cognizant of the compute capability of the NDP and cache locality on the host processor, system performance and energy are improved by up to 66.8% and 37.6%, respectively.

...read moreread less

57 citations

Posted Content•

Enabling the Adoption of Processing-in-Memory: Challenges, Mechanisms, Future Research Directions

[...]

Saugata Ghose, Kevin Hsieh, Amirali Boroumand, Rachata Ausavarungnirun, Onur Mutlu - Show less +1 more

01 Feb 2018-arXiv: Hardware Architecture

TL;DR: This work proposes and evaluates two general-purpose solutions that minimize unnecessary off-chip communication for PIM architectures and shows that both mechanisms improve the performance and energy consumption of many important memory-intensive applications.

...read moreread less

Abstract: Poor DRAM technology scaling over the course of many years has caused DRAM-based main memory to increasingly become a larger system bottleneck A major reason for the bottleneck is that data stored within DRAM must be moved across a pin-limited memory channel to the CPU before any computation can take place This requires a high latency and energy overhead, and the data often cannot benefit from caching in the CPU, making it difficult to amortize the overhead Modern 3D-stacked DRAM architectures include a logic layer, where compute logic can be integrated underneath multiple layers of DRAM cell arrays within the same chip Architects can take advantage of the logic layer to perform processing-in-memory (PIM), or near-data processing In a PIM architecture, the logic layer within DRAM has access to the high internal bandwidth available within 3D-stacked DRAM (which is much greater than the bandwidth available between DRAM and the CPU) Thus, PIM architectures can effectively free up valuable memory channel bandwidth while reducing system energy consumption A number of important issues arise when we add compute logic to DRAM In particular, the logic does not have low-latency access to common CPU structures that are essential for modern application execution, such as the virtual memory and cache coherence mechanisms To ease the widespread adoption of PIM, we ideally would like to maintain traditional virtual memory abstractions and the shared memory programming model This requires efficient mechanisms that can provide logic in DRAM with access to CPU structures without having to communicate frequently with the CPU To this end, we propose and evaluate two general-purpose solutions that minimize unnecessary off-chip communication for PIM architectures We show that both mechanisms improve the performance and energy consumption of many important memory-intensive applications

...read moreread less

56 citations

Cites background or methods from "A scalable processing-in-memory acc..."

..., [2, 52]) treat an entire application thread as a PIM kernel, in order to minimize the amount of synchronization and data sharing that takes place between the main CPU and main compute-capable memory....
[...]
...We believe ample future work potential exists on examining other solutions for these two challenges as well as our solutions for them within the context of other in-memory accelerators, such as those described in [2, 20, 68, 92, 93, 195, 196, 199, 200]....
[...]
...Several works [2, 3, 21, 68] provide a detailed explanation of this process....
[...]
...Ultimately, there is significant time and energy wasted on moving data between the CPU and memory, many times with little benefit in return, especially in workloads where caching is not very effective [2, 3]....
[...]
...Examples of 3D-stacked DRAM include High-Bandwidth Memory (HBM) [75, 115] and the Hybrid Memory Cube (HMC) [2, 71, 72]....
[...]

Proceedings Article•DOI•

Newton: A DRAM-maker’s Accelerator-in-Memory (AiM) Architecture for Machine Learning

[...]

Mingxuan He¹, Choung-Ki Song², Ilkon Kim², Chun-Seok Jeong², Seho Kim², Il Park², Mithuna Thottethodi¹, T. N. Vijaykumar¹ - Show less +4 more•Institutions (2)

Purdue University¹, SK Hynix²

01 Oct 2020

TL;DR: This work focuses on digital PIM which provides higher bandwidth than PNM and does not incur the reliability issues of analog PIM, and describes Newton, a major DRAM maker’s upcoming accelerator-in-memory (AiM) product for machine learning, which makes PIM feasible for the first time.

...read moreread less

Abstract: Advances in machine learning (ML) have ignited hardware innovations for efficient execution of the ML models many of which are memory-bound (e.g., long short-term memories, multi-level perceptrons, and recurrent neural networks). Specifically, inference using these ML models with small batches, as would be the case at the Cloud edge, has little reuse of the large filters and is deeply memory-bound. Simultaneously, processing-in or -near memory (PIM or PNM) is promising unprecedented high-bandwidth connection between compute and memory. Fortunately, the memory-bound ML models are a good fit for PIM. We focus on digital PIM which provides higher bandwidth than PNM and does not incur the reliability issues of analog PIM. Previous PIM and PNM approaches advocate full processor cores which do not conform to PIM’s severe area and power constraints. We describe Newton, a major DRAM maker’s upcoming accelerator-in-memory (AiM) product for machine learning, which makes the following contributions: (1) To satisfy PIM’s area constraints, Newton (a) places a minimal compute of only multiply-accumulate units and buffers in the DRAM which avoids the full-core area and power overheads of previous work and thus makes PIM feasible for the first time, and (b) employs a DRAM-like interface for the host to issue commands to the PIM compute. The PIM compute is rate-matched to the internal DRAM bandwidth and employs a non-intuitive, global input vector buffer shared by the entire channel to capture input reuse while amortizing buffer area cost. To the host, Newton’s interface is indistinguishable from regular DRAM without any offloading overheads and PIM/non-PIM mode switching, and with the same deterministic latencies even for floating-point commands. (2) To prevent the PIM-host interface from becoming a bottleneck, we include three optimizations: commands which gang multiple compute operations both within a bank and across banks; complex, multi-step compute commands – both of which save critical command bandwidth; and targeted reduction of t FAW overhead. (3) To capture output vector reuse with reasonable buffering, Newton employs an unusually-wide interleaved layout for the matrix. Our simulations running state-of-the-art neural networks show that building on a realistic HBM2E-like DRAM, Newton achieves 10x and 54x average speedup over a non-PIM system with infinite compute that perfectly uses the external DRAM bandwidth and a realistic GPU, respectively.

...read moreread less

55 citations

Proceedings Article•DOI•

Highly scalable near memory processing with migrating threads on the emu system architecture

[...]

Timothy J. Dysart, Peter M. Kogge, Martin Deneroff, Eric Bovell, Preston Briggs, Jay B. Brockman, Kenneth Jacobsen, Yujen Juan, Shannon K. Kuntz, Richard Lethin, Janice O. McMahon, Chandra Pawar, Martin Perrigo, Sarah Rucker, John Ruttenberg, Max Ruttenberg, Steve Stein - Show less +13 more

13 Nov 2016

TL;DR: A new, highly-scalable PGAS memory-centric system architecture where migrating threads travel to the data they access, and a comparison of key parameters with a variety of today's systems, of differing architectures, indicates the potential advantages.

...read moreread less

Abstract: There is growing evidence that current architectures do not well handle cache-unfriendly applications such as sparse math operations, data analytics, and graph algorithms. This is due, in part, to the irregular memory access patterns demonstrated by these applications, and in how remote memory accesses are handled. This paper introduces a new, highly-scalable PGAS memory-centric system architecture where migrating threads travel to the data they access. Scaling both memory capacities and the number of cores can be largely invisible to the programmer.The first implementation of this architecture, implemented with FPGAs, is discussed in detail. A comparison of key parameters with a variety of today's systems, of differing architectures, indicates the potential advantages. Early projections of performance against several well-documented kernels translate these advantages into comparative numbers. Future implementations of this architecture may expand the performance advantages by the application of current state of the art silicon technology.

...read moreread less

53 citations

Proceedings Article•DOI•

MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm

[...]

Wenqin Huangfu¹, Xueqi Li², Shuangchen Li¹, Xing Hu¹, Peng Gu¹, Yuan Xie¹ - Show less +2 more•Institutions (2)

University of California, Santa Barbara¹, Chinese Academy of Sciences²

12 Oct 2019

TL;DR: A practical, energy efficient, Dual-Inline Memory Module (DIMM) based, NDP Accelerator for DNA Seeding Algorithm (MEDAL), which is based on off-the-shelf DRAM components and an algorithm-specific data compression technique to reduce memory footprint, introduce more space for the data mapping, and reduce the communication overhead is proposed.

...read moreread less

Abstract: Computational genomics has proven its great potential to support precise and customized health care. However, with the wide adoption of the Next Generation Sequencing (NGS) technology, 'DNA Alignment', as the crucial step in computational genomics, is becoming more and more challenging due to the booming bio-data. Consequently, various hardware approaches have been explored to accelerate DNA seeding - the core and most time consuming step in DNA alignment. Most previous hardware approaches leverage multi-core, GPU, and FPGA to accelerate DNA seeding. However, DNA seeding is bounded by memory and above hardware approaches focus on computation. For this reason, Near Data Processing (NDP) is a better solution for DNA seeding. Unfortunately, existing NDP accelerators for DNA seeding face two grand challenges, i.e., fine-grained random memory access and scalability demand for booming bio-data. To address those challenges, we propose a practical, energy efficient, Dual-Inline Memory Module (DIMM) based, NDP Accelerator for DNA Seeding Algorithm (MEDAL), which is based on off-the-shelf DRAM components. For small databases that can be fitted within a single DRAM rank, we propose the intra-rank design, together with an algorithm-specific address mapping, bandwidth-aware data mapping, and Individual Chip Select (ICS) to address the challenge of fine-grained random memory access, improving parallelism and bandwidth utilization. Furthermore, to tackle the challenge of scalability for large databases, we propose three inter-rank designs (polling-based communication, interrupt-based communication, and Non-Volatile DIMM (NVDIMM)-based solution). In addition, we propose an algorithm-specific data compression technique to reduce memory footprint, introduce more space for the data mapping, and reduce the communication overhead. Experimental results show that for three proposed designs, on average, MEDAL can achieve 30.50x/8.37x/3.43x speedup and 289.91x/6.47x/2.89x energy reduction when compared with a 16-thread CPU baseline and two state-of-the-art NDP accelerators, respectively.

...read moreread less

53 citations

Cites background or methods from "A scalable processing-in-memory acc..."

..., the index of the first appearance of A,T ,C,G in sorted BR ; • OR [len + 1][4]: the occurrence array, i....
[...]
...• CR [4]: the accumulative count array, i....
[...]
...We expect applications such as graph processing [4], database searching [31], and sparse matrix computing [41] will also benefit from the proposed techniques....
[...]
...To perform the task described in Algorithm 1, the accelerator contains • registers to store the query sequence q, • a 4×64-bit register file to store CR [4], • a data reorganization engine to calculate OR [x] from its stored data structure, • two 64-bit unsigned adders to update Iuppper and I lower ,...
[...]
...1 Preprocess: Derive BR [len], SR [len], CR [4], and OR [len + 1][4]; 2 while I lower <= I do...
[...]

Collapse

A scalable processing-in-memory accelerator for parallel graph processing

Citations

Cites background or methods from "A scalable processing-in-memory acc..."

Cites background or methods from "A scalable processing-in-memory acc..."

References

"A scalable processing-in-memory acc..." refers methods in this paper

"A scalable processing-in-memory acc..." refers methods in this paper

"A scalable processing-in-memory acc..." refers methods in this paper

"A scalable processing-in-memory acc..." refers methods in this paper

Related Papers (5)