Data volume reduction and efficient data placement for accelerator-based HPC systems
Abstract
Details
- Title: Subtitle
- Data volume reduction and efficient data placement for accelerator-based HPC systems
- Creators
- Shihui Song
- Contributors
- Peng Jiang (Advisor)Steve Goddard (Committee Member)Mehrdad Moharrami (Committee Member)Tianyu Zhang (Committee Member)Robert Underwood (Committee Member)
- Resource Type
- Dissertation
- Degree Awarded
- Doctor of Philosophy (PhD), University of Iowa
- Degree in
- Computer Science
- Date degree season
- Spring 2026
- DOI
- 10.25820/etd.008441
- Publisher
- University of Iowa
- Number of pages
- xv, 145 pages
- Copyright
- Copyright 2026 Shihui Song
- Language
- English
- Date submitted
- 04/20/2026
- Description illustrations
- Illustrations, graphs, charts, tables
- Description bibliographic
- Includes bibliographical references (pages 130-145).
- Public Abstract (ETD)
As modern high-performance computing (HPC) and machine learning (ML) applications continue to grow in scale, data-centric bottlenecks increasingly dominate system performance. These bottlenecks include costly data movement as well as excessive memory and storage demands. This thesis addresses these challenges by developing architecture-conscious techniques for efficient data placement and data volume reduction, with the goal of enabling more scalable execution of big-data workloads. First, it studies data movement in large-scale graph neural network training and proposes an efficient data placement strategy that minimizes data loading time when moving graph features between CPU memory and multiple GPUs. It then presents CereSZ, the first end-to-end error-bounded lossy compression framework for the Cerebras CS-2 system, to help reduce the storage and communication costs of data-intensive applications. CereSZ establishes a compression pipeline with block-wise design, stage-wise pipelining, and coordinated execution across processing elements. Building on this effort, it introduces CereSZ-II, which further improves the computational efficiency and scalability of lossy compression on Cerebras through a more efficient fixed-size Huffman encoding method. Finally, it presents P3Z, a domain-specific compiler that automates the generation of high-performance lossy compression code for both CPU and Cerebras platforms, improving programmability while preserving performance.
Together, these contributions show that improving how data is moved, reduced, and programmed is essential for making future scientific and AI applications more scalable, efficient, and practical on modern computing platforms.
- Academic Unit
- Computer Science
- Record Identifier
- 9985176871802771