Shared matrix structure makes an implementation strategy reusable

Gated DeltaNet · scalar decay

Lij = βi γij (kiT kj)

4-token sketch

One shared decay factor across all key channels

M = I + L

Kimi Delta Attention · channel-wise decay

Lij = βi kiT diag(γij) kj

4-token sketch

A separate decay factor per key channel

M = I + L

k: normalized key; β: update weight; γij: decay accumulated over tokens j + 1, …, i. Only earlier tokens contribute: j < i.

Different coefficients; the same unit-lower-triangular solve M X = R.
Gated DeltaNet implementation Invert 8×8 diagonal blocks; merge into 16×16, 32×32, then 64×64.
Reusable strategy Block inversion uses the triangular structure shared by both systems.
Kimi Delta Attention implementation Reuse the merge steps with this operator's coefficients. Adapt block sizes and operand types.