Evaluating Processing-In-Memory (PIM) Topologies for Graph Processing Workloads
Author
Ms. Bhanupriya Sharma
Abstract
Large-scale graph processing algorithms (e.g., PageRank, Breadth-First Search, Single-Source Shortest Path) are foundational to web search, social network analysis, bioinformatics, and recommendation systems. However, conventional CPU/GPU compute architectures struggle to process graph algorithms efficiently. Graph workloads are inherently memory-bound: they exhibit low compute-to-memory ratios, irregular access patterns, poor spatial and temporal cache locality, and severe pointer-chasing behaviors. On traditional von Neumann systems, up to 85% of total execution cycles are spent stalling on off-chip DRAM interconnects, leading to substantial energy waste.
Processing-In-Memory (PIM) has emerged as a promising architectural paradigm to break the 'memory wall' by shifting compute primitives directly inside or adjacent to memory structures. However, PIM is not a monolith: hardware implementations range from fine-grained bank-level PIM (e.g., UPMEM DPU, Samsung FIMDRAM) to 3D-stacked logic-layer PIM (e.g., HBM/HMCC-style) and CXL-attached disaggregated near-memory accelerators. This paper presents a cycle-accurate architectural evaluation comparing three representative PIM topologies against a high-end server-grade multi-socket Xeon processor across four canonical graph benchmarks (BFS, PageRank, SSSP, Connected Components) operating on real-world and synthetic scale-free graphs (up to Scale 28).
Using a gem5-based cycle-accurate simulation framework integrated with Ramulator2, our results demonstrate that 3D-stacked logic-layer PIM achieves an average 8.4x speedup and a 7.2x reduction in energy consumption over baseline CPU architectures by leveraging ultra-wide TSV bandwidth. Bank-level PIM achieves superior memory bandwidth (14.2x) but suffers from strict bank-to-bank inter-processor communication overheads during global graph traversal phases. We synthesize these architectural trade-offs into a quantitative design guideline matrix for next-generation graph processing hardware.
Keywords
While Bank-Level PIM offers massive raw internal memory bandwidth (up to 1.8 TB/s), 3D-Stacked Logic-Layer PIM provides the optimal balance for irregular graph traversals due to its flexible inter-stack interconnect
Full Text:
Download Paper PDF
References
- Ahn, J., et al. (2015). A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing. ISCA '15.
- Ke, L., et al. (2020). Near-Memory Processing in the Era of Big Data: A Survey and Architecture Classification. IEEE Micro.
- Devaux, A., et al. (2021). Evaluated Performance of the UPMEM Processing-in-Memory System. IEEE Access.
- Kwon, Y., et al. (2021). 25.4 A 20nm 6GB Function-In-Memory DRAM (FIMDRAM) for AI Acceleration. ISSCC '21.
- Boroumand, A., et al. (2018). Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks. ASPLOS '18.
Share your valuable work from Social Media Buttons