← All writing
GPUsCUDA

Learning CUDA with PMPP, Part 2: Heterogeneous Data-Parallel Computing

Data parallelism reorganizes computation around independent pieces of data, while task parallelism isolates independent portions of an application.

Data parallelism refers to a computing phenomenon where computations to be performed on a dataset are independent of one another and can therefore be performed in parallel.

Writing a data-parallel program often entails reorganizing the computation around the data so that the computations can be independent of one another.

Color-to-grayscale conversion

Color-to-grayscale conversion is a great example of data parallelism. Every image is made up of pixels containing RGB (red, green, and blue) values ranging from 0 (dark) to 1 (full intensity). To compute the grayscale luminance L of a pixel, we apply the weighted-sum formula:

L = r × 0.299 + g × 0.587 + b × 0.114
  • r, g, and b are the RGB values of that pixel.
  • Every individual pixel can be converted to grayscale independently of the others.

Task parallelism

Aside from parallelism in data, we can also exploit parallelism in computations. This is referred to as task parallelism. Through task decomposition, we can isolate portions of an application that can be executed independently of one another.

CUDA C++ program structure

These terms define the CUDA C++ execution model:

Grid
A grid is a collection of threads. Threads in a grid execute a kernel function and are divided into thread blocks.
Thread block
A thread block is a group of threads that execute on the same streaming multiprocessor (SM). Threads within a thread block have access to shared memory and can be explicitly synchronized.
Kernel function
A kernel function is an implicitly parallel subroutine that executes under the CUDA execution and memory model for every thread in a grid.
Host
The host is the execution environment that initially invoked CUDA—typically the thread running on a system’s CPU.
Parent
A parent thread, thread block, or grid is one that has launched one or more child grids. The parent is not considered complete until all of its child grids have completed.
Child
A child thread, block, or grid is one launched by a parent grid. A child grid must complete before its parent thread, thread block, or grid is considered complete.
Thread-block scope
Objects with thread-block scope have the lifetime of a single thread block. Their behavior is defined only when operated on by threads in the block that created them, and they are destroyed when that block completes.
Device runtime
The device runtime is the runtime system and set of APIs that enable kernel functions to use dynamic parallelism.