From TIRx-lite source to generated code
Schematic. Click a source line to see the code it becomes.
TIRx-lite source
Generated CUDA
__global__ void vector_add_kernel(...) {
int i = blockIdx.x * 128 + threadIdx.x;
if (i < n) {
asm("ld.global.f32 %0, [%1];" ...);
asm("add.f32 %0, %1, %2;" ...);
asm("st.global.f32 [%0], %1;" ...);
Kernel entry
@txl.kernel fixes the CTA width and becomes a CUDA kernel.
Coordinates
The entry's block and thread indices become blockIdx and threadIdx.
Control flow
txl.If is CUDA-level control flow and becomes an ordinary if.
One call, one instruction
The load becomes exactly one ld.global.f32.
One call, one instruction
The registration of add.f32 declares which operands it reads and writes and supplies the PTX template.
One call, one instruction
The store becomes exactly one st.global.f32.