CUDAAIComputing1 min read

Introduction to CUDA

A GPU was first built to draw graphics. To draw a picture, it has to calculate the color of millions of pixels, and it can do many of these small calculations at the same time. CUDA is a platform created by NVIDIA that lets us use this power for our own programs, not only for graphics. It works with NVIDIA GPUs. In this content, we will see the main idea behind it without installing anything: how a big job is split into many tiny pieces.

A CPU has a small number of powerful cores. A GPU has thousands of simpler ones. Imagine we need to paint every pixel of an image. The CPU paints a few pixels at a time, but very quickly. The GPU paints many pixels at once, even though each of its cores is slower. Run both of them, then try the smallest size.

Interactive
CPU4 fast cores
0.0 ticks
GPU256 simple cores
0.0 ticks

A simplified model, not real measurements. Real GPUs have hundreds to thousands of cores, and real timings depend on the hardware.

When the job is small, the CPU wins. Before the GPU can start, the data usually has to be copied to its own memory, and the work has to be started. With only a few pixels, this setup takes longer than the work itself. As the job grows, the GPU leaves the CPU far behind. That is why GPUs are great for work that can be split into many independent pieces, like images and the math behind neural networks.

In CUDA, the smallest worker is a thread. Threads are grouped into blocks, and all the blocks together are called a grid. When we start work on the GPU, we choose how many blocks we want and how many threads each block has.

Every thread runs the same code. So how does a thread know which piece of the work is its own? CUDA gives each thread a few built-in values. blockIdx.x is the number of its block, blockDim.x is how many threads a block has and threadIdx.x is its number inside the block. With a small formula, we turn them into one unique number for every thread. Hover over the threads below to see it. When it feels clear, press Quiz me and try to find the right thread yourself.

Interactive
Grid · gridDim.x = 3
block 0
block 1
block 2
int i = blockIdx.x * blockDim.x + threadIdx.x;
// i = 1 * 4 + 2 = 6
Array · 12 elements, one per thread
0
1
2
3
4
5
6
7
8
9
10
11

You may wonder why every name ends with .x. Blocks and threads can also be arranged in two or three dimensions, with .y and .z, which is handy for images and 3D data. Here we only use one dimension to keep things simple.

A function that we launch on the GPU like this is called a kernel. We mark it with __global__ and start it with the special <<<blocks, threads>>> syntax. Below is the classic first CUDA program: adding two arrays. Each thread adds only one pair of numbers, and the threads work in parallel.

There is a catch. Threads come in whole blocks, so we often start a few more threads than we need. If n is 10 and a block has 4 threads, we need 3 blocks, which means 12 threads. That is why each thread checks i < n before it writes. Real programs usually use bigger blocks, like 256 threads, but small numbers are easier to see. Change the number of elements and try turning off the bounds check.

Interactive
Kernel
__global__ void add(float *a, float *b, float *c, int n) {
int i = blockIdx.x * blockDim.x + threadIdx.x;
if (i < n)
c[i] = a[i] + b[i];
}
 
// launch
add<<<3, 4>>>(a, b, c, 10);
thread
0
1
2
3
0
1
2
3
0
1
2
3
a
1
4
7
3
6
2
5
1
4
7
b
1
6
2
7
3
8
4
9
5
1
c
3 blocks of 4 threads = 12 threads for 10 elements. Press Run.

On most computers, the CPU and the GPU have separate memories. Before a kernel runs, the data is copied to the GPU, for example with cudaMemcpy, and the result is copied back at the end. These copies can take longer than the work itself. If we copy the data back and forth around every kernel, the GPU can spend most of its time waiting. If we keep the data on the GPU between kernels, we only pay for the copies once.

Interactive
CPU memory
RAM
GPU memory
VRAM
Press Run to watch the data move.
Copy around every kernel21 ticks
→
←
→
←
→
←
Keep data on the GPU9 ticks
→
←

14% of the time is real work, 86% is moving data around.

copy kernel

A simplified model: a copy takes 3 ticks and a kernel takes 1. Real numbers depend on the data size and the hardware.

You meet this even if you never write CUDA yourself. In PyTorch, .to("cuda") copies a tensor into GPU memory, and .cpu() or .item() bring data back to the CPU. .item() also makes the CPU wait until the GPU has finished its work. That is why calling them again and again inside a training loop can slow it down: the data keeps moving back and forth, and the CPU keeps waiting.

Now we know all the pieces, so let's put them together. A small CUDA program has five steps: reserve memory on the GPU, copy the inputs, run the kernel, copy the result back and free the memory. Step through it below and watch both memories.

Interactive
main
const int n = 10;
size_t size = n * sizeof(float);
float a[n], b[n], c[n]; // on the CPU, a and b are filled
float *d_a, *d_b, *d_c; // will point to GPU memory
 
cudaMalloc(&d_a, size);
cudaMalloc(&d_b, size);
cudaMalloc(&d_c, size);
cudaMemcpy(d_a, a, size, cudaMemcpyHostToDevice);
cudaMemcpy(d_b, b, size, cudaMemcpyHostToDevice);
int threads = 256;
int blocks = (n + threads - 1) / threads;
add<<<blocks, threads>>>(d_a, d_b, d_c, n);
cudaMemcpy(c, d_c, size, cudaMemcpyDeviceToHost);
cudaFree(d_a); cudaFree(d_b); cudaFree(d_c);
cudaMalloc gives us empty space in GPU memory. d_ is a common prefix for GPU (device) data.
CPU memory
a
1
4
7
3
6
2
5
1
4
7
b
1
6
2
7
3
8
4
9
5
1
c
2
10
9
10
9
10
9
10
9
8
GPU memory
d_a
1
4
7
3
6
2
5
1
4
7
d_b
1
6
2
7
3
8
4
9
5
1
d_c
2
10
9
10
9
10
9
10
9
8

That is the main idea of CUDA. Split the work into many small pieces, let every thread find its own piece with a simple formula, and move the data as little as possible. Many GPU programs, from image filters to neural networks, are built on these same ideas.


If this content helped you, you can buy me a coffee.
RelatedStable Diffusion Prompt Guide

You can join the newsletter to be notified of awesome interactive articles and courses about software, design and AI. You will receive at most a few emails per month.


Join our Supporters