GPU Programming Projects
Write real GPU kernels in Triton and Numba CUDA: fused softmax, tiled matmul, image filters and reductions.
4 projects · GPU notebook · Premium only
Image-Processing Kernels with Numba CUDA
PROWrite your own CUDA kernels in Python with Numba and run them on a real 4.5-megapixel photo: RGB-to-grayscale with one thread per pixel, a 9x9 box blur (first straight from global memory, then tiled through shared memory with a halo), and a Sobel edge detector. Every kernel is checked against NumPy/SciPy, then benchmarked properly (warm-up, synchronize) to see where the GPU wins. The photo (~5 MB) is downloaded from the Hugging Face Hub on first run.
Parallel Reductions & Histograms with Numba CUDA
PROMany threads writing to the same place is the classic GPU problem. Sum 16.8 million floats twice, first with one atomic add per element and then with a shared-memory tree reduction, and build a 256-bin histogram of an 18-megapixel photo with global atomics and then with a privatized per-block histogram. Every result is checked against NumPy, and every kernel is timed and compared in GB/s. The photo (~5 MB) is downloaded from the Hugging Face Hub on first run.
Write a Fused Softmax Kernel in Triton
PROWrite your first GPU kernels in Triton. You'll warm up with a masked vector-add kernel, then build a fused, numerically stable row-wise softmax that reads each row once, keeps it in registers and writes it once. You'll check it against torch.softmax on awkward shapes and extreme logits, then benchmark it with triton.testing.do_bench against torch.softmax and a naive multi-kernel PyTorch version, plotting effective GB/s. Nothing to download.
Tiled Matrix Multiplication in Triton
PROBuild a matrix multiply in Triton from scratch. You'll write a blocked fp16 kernel that walks the K dimension in BLOCK_K steps with tl.dot and an fp32 accumulator, masking ragged edges, and check it against torch.matmul on shapes that aren't multiples of the tile. You'll read the compiled PTX to see what tl.dot became: on the T4 (Turing, sm_75) Triton 3 runs it as FMAs on the CUDA cores, not on the tensor cores. Then you'll add grouped tile ordering, @triton.autotune over a small set of tile configs and a fused leaky-ReLU epilogue, and benchmark TFLOPS against cuBLAS in fp16 and fp32. The goal is to learn the mechanics of tiling and autotuning, not to beat cuBLAS. Nothing to download.