[Submitted on 24 Jul 2026 (v1), last revised 28 Jul 2026 (this version, v2)]
Abstract:We present $g$MAGNUS, a novel algorithm for sparse matrix-matrix multiplication (SpGEMM) of irregular matrices on GPUs. Such matrices often contain many heavy rows, those with large intermediate products that force local memory accumulators to spill to global memory. $g$MAGNUS addresses this by computing an intra-row reordering of intermediate products, subdividing heavy rows into independent chunks that can be accumulated completely in local memory. This reordering uses novel outer product and hierarchical multisplit operations. The algorithm is input- and system-aware, automatically determining the number of chunks and multisplit levels based on the input matrix dimensions and local memory size. Experimental results on two extensive datasets show that $g$MAGNUS achieves a geometric-mean speedup of 1.81 to 7.62$\times$ over five leading algorithms (including MKL and cuSPARSE) on Intel Ponte Vecchio and NVIDIA H200. Additionally, the core kernels of $g$MAGNUS are evaluated, achieving near-peak performance compared to their theoretical upper bound.
Submission history
From: Jordi Wolfson-Pou [view email]
[v1]
Fri, 24 Jul 2026 19:11:46 UTC (438 KB)
[v2]
Tue, 28 Jul 2026 02:28:40 UTC (438 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.