Chapter 2: Moving Data: GPU Memory with Numba
The GPU has its own memory, separate from your program's. Learn how to move arrays to the device and back, allocate and mutate device arrays, index 2D grids for matrices, and use fast per-block shared memory with a barrier to let threads cooperate.
Chapter Overview
A kernel can only touch memory that lives on the GPU. Your NumPy arrays live in host (CPU) memory, in a completely separate address space connected to the GPU by the PCIe bus. Before a kernel can run, its inputs must be copied to the device, and after it finishes, the results must be copied back.
Managing that movement deliberately is most of what separates a fast GPU program from a slow one — transfers are expensive, so you want to move data up once, do as much work as possible on the device, and move only the results back.
This chapter covers the memory model and the tools Numba gives you to work with it:
- Host vs Device Memory: two separate worlds, bridged by explicit copies
- to_device & copy_to_host: staging inputs and retrieving results
- 2D Grids for Matrices: indexing rows and columns with cuda.grid(2)
- Shared Memory & __syncthreads(): fast on-chip memory that lets a block's threads cooperate
Chapter Roadmap
Click any topic to jump in
Host vs Device Memory
CPU and GPU have separate address spaces bridged by PCIe; transfers are explicit and costly.
to_device & copy_to_host
Stage inputs with to_device, allocate outputs with device_array, retrieve with copy_to_host.
2D Grids for Matrices
cuda.grid(2) gives each thread a (row, col); 2D blocks tile the matrix; guard both dimensions.
Shared Memory & __syncthreads()
Fast per-block on-chip memory; a barrier coordinates cooperative load-then-read across threads.
Sign up to unlock this chapter
This chapter is part of PixelBank Premium. Create a free account, then upgrade to read the full lesson — concepts, walkthroughs, and exercises.