NVIDIA's Gated DeltaNet-2 — linear attention with separate erase and write gates
Linear attention reduces unbounded KV cache to fixed recurrent state, but updating memory without corrupting past associations stays tricky. NVIDIA decoupled the single gate into channel-wise erase (key axis) and write (value axis) controls.
At 1.3B params on 100B tokens, Gated DeltaNet-2 beats Mamba-2, Mamba-3, KDA, and the original Gated DeltaNet on language modeling, reasoning, and long-context retrieval—strongest on RULER and multi-key needle tasks.