One kernel, three views

The same addition at each level of vector_add

TIRx-lite source
txl.ptx.add.f32(x, x, y)
vector_add
TIRx IR
T.ptx.add(
    buffer, buffer,
    buffer_1, ...)
vector_add.func,
read by the analyses
Generated CUDA
asm("add.f32 %0, %1, %2;"
    : "=f"(__d)
    : "f"(__a), "f"(__b));
vector_add.source()
Output abridged.
CUDA → ptxas → Machine code