Each warp has a clock [W3,W7]. Its own entry counts its modeled events; the other entry records the latest event ordered before it from the other warp. Both start at after common setup.
Pseudocode · one CTA
Each warpgroup has four warps. Setup synchronization precedes the TMEM accesses; cleanup synchronization follows them.
tilerows , columns 0..31Clocks at the conflicting warp instructions
Warp 3 · group 0
Warp 7 · group 1
Read each clock as an ordering history
The store's clock records W3's first event, with no preceding W7 event. The load's clock records W7's first event, with no preceding W3 event.
The zero in the load's W3 entry means the store is missing from the load's ordering history. Even if W3 runs first on the GPU, no synchronization orders its store before W7's load.
The clocks merge at cleanup, after both accesses
The local waits advance only their own warp's entry. Cleanup then merges both histories: taking the maximum keeps the latest event from each warp. Each warp also increments its own entry for cleanup.
max(, ) = W3 → ; W7 → The recorded store and load clocks remain and . This later synchronization does not order those earlier accesses.
Compare the clocks recorded at lines 4 and 7
Store S = ; load L = . S ≤ L would mean that the load's history includes the store. It requires every entry of S to be no larger than the corresponding entry of L.
| Comparison | W3 | W7 |
|---|---|---|
| S ≤ L? | ||
| L ≤ S? |
Neither access happens before the other. They conflict on the same TMEM locations: a data race.