GPU Programming with CUDA
NVIDIA introduced CUDA in 2007, providing developers with a general-purpose programming model to leverage GPU computational power. This article examines GPU programming models using NVIDIA GPU architecture as reference.
SIMD vs SIMT Execution Models
SIMD (Single Instruction Multiple Data) executes a single instruction stream with multiple data inputs processed simultaneously. Most AI chips employ this hardware architecture. Vector addition in SIMD appears as:
[VLOAD, VLOAD, VADD, VSTORE], VECTOR_LENGTH
SIMT (Single Instruction Multiple Threads) executes scalar instructions across multiple thread streams, dynamically grouping threads into warps for execution. Vector addition in SIMT appears as:
[LOAD, LOAD, ADD, STORE], THREAD_COUNT
NVIDIA GPUs utilize SIMT execution, offering several advantages:
- Eliminates developer effort to organize data into specific vecter lengths
- Hardware-level resolution of SIMD data path pipeline scheduling
- Independent thread execution enabling flexible branching
- Dynamic warp organization of threads executing identical instructions
Consider a warp containing 32 threads processing 32,000 iterations. This requires 1,000 warps, with Warp0 handling threads 0-31, Warp1 handling threads 32-63, and Warp20 processing threads 661-692. SIMT enables dynamic thread grouping that increases parallelism.
Warps and Fine-Grained Multithreading
SIMT architecture employs Fine-Grained Multi-Threading (FGMT) to subdivide processor pipelines into smaller units, enabling interleaved instruction execution across threads. This reduces latency and resource waste while achieving memory-computation parallelism.
A warp represents a thread collection executing identical instructions across different memory addresses. Multiple warps compose SIMD pipelines for operations. FGMT hides latency through out-of-order warp execution rather than sequential scheduling. Thread register values remain in Register Files, with FGMT accommodating long latencies.
NVIDIA's hardware warp scheduler enables memory operations to complete before instruction execution in SIMD pipelines, hiding memory access times across warps. Each warp contains 32 threads and 8 execution lanes, performing 24 operations per warp execution cycle.
At the macroarchitecture level, GPU data flows from GDDR through memory controllers to interconnection networks, then distributes to execution cores (CUDA Core/Tensor Core) containing SIMD execution units.
AMD Programming Model
AMD GPUs contain numerous compute units and cores but employ different programming approaches. In 2016, AMD launched Radeon Open Computing platform (ROCm), providing software support for Radeon GPU hardware, mathematical libraries, and modern programming languages. ROCm substantially supports CUDA compatibility while establishing an alternative ecosystem.
Before CDNA architecture debuted in 2020, AMD data center GPUs used GCN architecture. While successful in gaming consoles, GCN achieved limited data center market penetration due to performance limitations. AMD now separates GPU architecture development into CDNA (computation) and RDNA (graphics) lines.
AMD's MI300X accelerator employs accelerator complex dies (XCD), each containing core sets and shared cache. Each XCD physically contains 40 CDNA 3 compute units, with 38 enabled per XCD in MI300X. XCDs include 4MB L2 cache serving all compute units. MI300X incorporates 8 XCDs totaling 304 compute units presented as a single GPU.
Each compute unit contains 4 SIMD execution units. Per cycle, CU schedulers select one SIMD for execution while verifying thread readiness. MI300 supports ROCm 6 with TF32/FP8 data types, Transformer Engine, structured sparsity, and AI/ML frameworks.
Comparatively, NVIDIA's H100 contains 132 streaming multiprocessors (SMs) presented as a unified GPU. Computation distributes via CUDA programs to execution cores containing SIMD units.