Learning CUDA with PMPP, Part 1: Introduction to Parallel Programming
As performance gains from newer chips plateau, developers need to understand parallel programming and the different design goals of CPUs and GPUs.
Applications are demanding more and more compute power as Moore’s law has stopped applying. The speedup developers saw from simply running code on newer chips has begun to plateau.
Most programs were written and executed sequentially. A thread is informally a sequence of instruction-execution activities resulting from the sequential, stepwise execution of an application.
The concurrency revolution helped pave the way for improving the performance of modern programs, as essentially all microprocessors today are parallel computers.
This means there is a greater need for developers to learn about parallel programming, which is the focus of Programming Massively Parallel Processors.
Multicore versus many-thread
The multicore paradigm seeks to preserve the fast execution speed of sequential programs while spreading execution across multiple cores.
The many-thread paradigm focuses on improving the throughput of parallel execution by spreading execution across a massive number of threads.

CPU and GPU design goals
The CPU is optimized for sequential-code performance. Its arithmetic logic units (ALUs) are designed to minimize the effective latency of operations. The tradeoff is greater chip-area and power usage. CPUs also have large L1 caches to replace long-latency memory access with short-latency cache access.
The GPU is designed and optimized for throughput. Video games and deep learning tasks require a massive number of floating-point calculations per second.
GPUs must also support loading and offloading large amounts of data in their DRAM. They achieve this with high memory bandwidth.
Increasing throughput is much cheaper than reducing latency in terms of chip area and power consumption. GPUs are therefore optimized to increase the throughput of the card rather than reduce the time required to execute individual threads.

In 2007, NVIDIA released CUDA, which represented a change in how GPUs were programmed. CUDA enabled a shift away from fixed-function programming toward general-purpose GPU programming. It extended C and C++ and democratized general-purpose programming on GPUs.
Amdahl’s law
Amdahl’s law states that the achievable performance improvement of a parallelized application is strictly limited by its sequential components. Regardless of how much the parallelizable segments are optimized, the non-parallelizable fraction imposes a definitive ceiling on total speedup.
The mathematical expression for Amdahl’s law is:
- f
- The portion of the application capable of parallel execution.
- 1 − f
- The sequential or serial fraction of the program.
- N
- The number of processors or threads used for parallel tasks.

Why parallel programming is difficult
Parallel programming is hard if you care about performance. It is challenging to write parallel algorithms with the same computational or algorithmic complexity as sequential algorithms. Often, non-intuitive techniques are needed to solve certain problems.
Work efficiency and algorithmic primitives such as prefix sums and scans are concepts I will learn through this book.
Achieving speedup in memory-bound applications—applications limited by memory-access speed—also requires methods for improving memory access.
In addition, the execution speed of parallel programs is very sensitive to input data. Data can vary in its characteristics, including its size and distribution. PMPP introduces techniques for regularizing data distributions and tailoring the assignment of data to threads.
Some parallel applications require little collaboration between threads and are often described as embarrassingly parallel. Other applications require significant collaboration between threads and synchronization operations such as barriers and atomic operations. This introduces overhead as threads wait for one another to complete. PMPP provides strategies for reducing this overhead.
The parallel programming interface used in PMPP is CUDA. However, the techniques covered are not limited to CUDA. Other parallel interfaces, such as OpenMP and MPI, will be more accessible with the techniques learned in this book.
Goals of the book
- Teach the reader how to program massively parallel processors to achieve high performance.
- Teach parallel programming for correct functionality and reliability.
- Future-proof parallel programs so they achieve speedup when ported to future hardware.
Fundamental concepts
- CUDA programming model: kernels, threads, blocks, and grids.
- GPU execution architecture: SMs, warps, scheduling, and SIMD.
- Memory hierarchy: registers, shared memory, caches, and global memory.
- Performance fundamentals: occupancy, divergence, coalescing, latency hiding, and tiling.
Parallel patterns
- Convolution and stencil.
- Histogram.
- Reduction.
- Prefix sum and scan.
- Merge and sort.
- Highly optimized matrix multiplication.
Advanced patterns and applications
- Dynamic programming and wavefront algorithms.
- Sparse matrices.
- Graph algorithms.
- Convolutional neural networks.
- Large language models, including batching and KV caching.
- Scientific computing.
- Multi-GPU and cluster programming.