Write GPU kernels by hand — threads, blocks, shared memory, and atomics with Numba, then OpenAI Triton for fused activations, reductions, softmax, matmul, and attention.