Shared matrix structure makes an implementation strategy reusable
Gated DeltaNet · scalar decay
Lij = βi γij (kiT kj)
One shared decay factor across all key channels
M = I + L
Kimi Delta Attention · channel-wise decay
Lij = βi kiT diag(γij) kj
A separate decay factor per key channel
M = I + L
k: normalized key; β: update weight; γij: decay accumulated over tokens j + 1, …, i. Only earlier tokens contribute: j < i.
Different coefficients; the same unit-lower-triangular solve M X = R.
Gated DeltaNet implementation
Invert 8×8 diagonal blocks; merge into 16×16, 32×32, then 64×64.
Reusable strategy
Block inversion uses the triangular structure shared by both systems.
Kimi Delta Attention implementation
Reuse the merge steps with this operator's coefficients.
Adapt block sizes and operand types.