📘
Triton Masked Copy Kernel
Problem Statement
Copy a 1D tensor into an output buffer using a Triton kernel, correctly handling a length that is not a multiple of BLOCK_SIZE.
Background
The final block usually has lanes past the end of the tensor. mask = offsets < n ensures those lanes are not loaded or stored, preventing out-of-bounds memory access.
Your Task
Implement copy_kernel and run(n=1000) (deliberately non-power-of-two) that returns whether the copy is exact.
How it is tested
Your solution must define a top-level function run(...) that allocates inputs on the GPU, launches your Triton kernel, and returns a boolean from torch.allclose(triton_out, torch_reference, ...). The grader prints run(...); the expected output is True.
Example:
Input:
n = 1000
Output:
True
Reasoning:
- The input
n = 1000is used to allocate a 1D tensor of length 1000 on the GPU. - The
copy_kernelfunction is launched with this tensor, using aBLOCK_SIZEthat does not divide 1000 evenly, resulting in a final block with lanes past the end of the tensor. - To prevent out-of-bounds memory access, a mask is applied using
mask = offsets < n, ensuring that only valid lanes are loaded and stored. - The copied tensor is then compared to a PyTorch reference implementation using
torch.allclose, which checks for element-wise equality within a tolerance, resulting in the outputTrueif the copy is exact.
Constraints:
- Must use a mask so the tail block does not read/write OOB
- out must equal x exactly
Editor
Python 3.13.1
GPU · T4
Test Results
0/0Run code to see test results.