Logo image
Data-efficient and fault-tolerant exascale computing
Dissertation   Open access

Data-efficient and fault-tolerant exascale computing

Yafan Huang
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Spring 2026
DOI: 10.25820/etd.008442
pdf
Yafan_Huang___University_of_Iowa_PhD_Thesis11.48 MBDownloadView
Open Access

Abstract

The emergence of exascale high-performance computing (HPC) systems has enabled unprecedented computational capability, driven by massive parallelism and heterogeneous architectures. However, this rapid growth has fundamentally intensified two critical challenges: managing the explosive increase in data volume and ensuring reliable execution on increasingly complex and error-prone hardware platforms. These trends make data efficiency and fault tolerance critical system concerns that directly impact performance, scalability, and scientific productivity. This dissertation aims to enable data-efficient and fault-tolerant exascale computing through efficient, practical, and flexible system software techniques. It focuses on two key research thrusts: high-performance data compression for heterogeneous accelerators with critical performance requirements and software-directed fault tolerance for large-scale HPC and AI workloads. On the data-efficient computing side, this dissertation presents a family of ultra-fast GPU lossy compression frameworks that achieve both high compression ratios and high runtime performance. It introduces a kernel-fused design that eliminates CPU involvement and reduces global synchronization overhead, demonstrating high end-to-end throughput, which is also the key metric for inline compression. It further develops co-optimization techniques across algorithm design and GPU execution, including efficient outlier encoding, vectorized memory access, and compressionaware prefix-sum, enabling simultaneous improvements in throughput and compression ratio. Finally, it proposes a general-purpose compression framework that supports dimension-aware processing, multi-algorithm support, memory-efficient compression, and selective decompression, significantly broadening applicability across scientific simulations and AI workloads. On the fault-tolerant computing side, this dissertation develops software-directed soft error detection techniques that improve detection capability, performance efficiency, and practical usability. It introduces MinpSID, a selective instruction duplication approach that enhances robustness across varying program inputs through combined static and dynamic analysis. It further presents ConDa, a compiler-optimized co-design that integrates instruction duplication with control-flow encoding to achieve comprehensive datapath protection with reduced overhead. Bev yond traditional HPC applications, this dissertation also proposes a versatile fault injector for large language model inference, investigates error propagation during this process, and presents several key takeaways and lightweight mitigation strategies. To demonstrate real-world impact, the proposed techniques have been integrated into domain science workflows, including reverse time migration for seismic imaging in collaboration with Saudi Aramco and data reduction pipelines for advanced light source experiments at Argonne National Laboratory. These case studies show that the proposed methods can significantly reduce data movement costs and memory footprint consumption while preserving scientific fidelity and enabling scalable execution in production environments. Collectively, this dissertation establishes a systematic approach to designing data reduction and fault tolerance techniques for modern heterogeneous exascale computing systems. It demonstrates that through careful co-design across algorithms, compilers, and hardware-aware execution, it is possible to simultaneously achieve high performance, high efficiency, and practical deployability for exascale scientific and AI computing.
Data Compression Compiler Optimization Fault Tolerance High-performance Computing Machine Learning Systems Parallel Computing Electrical engineering

Details

Metrics

1 Record Views
Logo image