Each warp has a clock [W3,W7]. Its own entry counts its modeled events; the other entry records the latest event ordered before it from the other warp. Both start at after common setup.

Pseudocode · one CTA

1group = warp // 4 2row = (warp % 4) * 32 + lane 3if group == 0: 4 tmem_store(tile, row, values) 5 wait_own_stores() 6if group == 1: 7 values = tmem_load(tile, row) 8 wait_own_loads() 9 output[row, :] = values

Each warpgroup has four warps. Setup synchronization precedes the TMEM accesses; cleanup synchronization follows them.

Both warps access tile
rows , columns 0..31

Clocks at the conflicting warp instructions

Warp 3 · group 0

Store · line 4
Wait for own storesline 5

Warp 7 · group 1

Load · line 7
Wait for own loadsline 8
After cleanup synchronization

Read each clock as an ordering history

The store's clock records W3's first event, with no preceding W7 event. The load's clock records W7's first event, with no preceding W3 event.

The zero in the load's W3 entry means the store is missing from the load's ordering history. Even if W3 runs first on the GPU, no synchronization orders its store before W7's load.