PIXELBANKv9.1.0
Menu
All topics

GPU Programming Projects

Write real GPU kernels in Triton and Numba CUDA: fused softmax, tiled matmul, image filters and reductions.

4 projects · GPU notebook · Premium only

Difficulty

Image-Processing Kernels with Numba CUDA

PRO

Write your own CUDA kernels in Python with Numba and run them on a real 4.5-megapixel photo: RGB-to-grayscale with one thread per pixel, a 9x9 box blur (first straight from global memory, then tiled through shared memory with a halo), and a Sobel edge detector. Every kernel is checked against NumPy/SciPy, then benchmarked properly (warm-up, synchronize) to see where the GPU wins. The photo (~5 MB) is downloaded from the Hugging Face Hub on first run.

GPU ProgrammingMedium bee.jpg (Hugging Face documentation-images) 4 sections 8 cells
GPU notebook

Parallel Reductions & Histograms with Numba CUDA

PRO

Many threads writing to the same place is the classic GPU problem. Sum 16.8 million floats twice, first with one atomic add per element and then with a shared-memory tree reduction, and build a 256-bin histogram of an 18-megapixel photo with global atomics and then with a privatized per-block histogram. Every result is checked against NumPy, and every kernel is timed and compared in GB/s. The photo (~5 MB) is downloaded from the Hugging Face Hub on first run.

GPU ProgrammingMedium 16.8M random floats + bee.jpg (Hugging Face documentation-images) 4 sections 9 cells
GPU notebook

Write a Fused Softmax Kernel in Triton

PRO

Write your first GPU kernels in Triton. You'll warm up with a masked vector-add kernel, then build a fused, numerically stable row-wise softmax that reads each row once, keeps it in registers and writes it once. You'll check it against torch.softmax on awkward shapes and extreme logits, then benchmark it with triton.testing.do_bench against torch.softmax and a naive multi-kernel PyTorch version, plotting effective GB/s. Nothing to download.

GPU ProgrammingMedium Synthetic tensors 4 sections 7 cells
GPU notebook

Tiled Matrix Multiplication in Triton

PRO

Build a matrix multiply in Triton from scratch. You'll write a blocked fp16 kernel that walks the K dimension in BLOCK_K steps with tl.dot and an fp32 accumulator, masking ragged edges, and check it against torch.matmul on shapes that aren't multiples of the tile. You'll read the compiled PTX to see what tl.dot became: on the T4 (Turing, sm_75) Triton 3 runs it as FMAs on the CUDA cores, not on the tensor cores. Then you'll add grouped tile ordering, @triton.autotune over a small set of tile configs and a fused leaky-ReLU epilogue, and benchmark TFLOPS against cuBLAS in fp16 and fp32. The goal is to learn the mechanics of tiling and autotuning, not to beat cuBLAS. Nothing to download.

GPU ProgrammingHard Synthetic fp16 matrices 4 sections 9 cells
GPU notebook