flash-linear-attention
library Your tags
Your notes
The de-facto reference Triton kernel library for emerging linear-attention and SSM architectures (GLA, DeltaNet, Mamba-2, RWKV, Gated DeltaNet; ~5.4K★), led by Songlin Yang (MIT, advised by Yoon Kim) with a multi-institution maintainer community, plus the companion flame training framework. The underlying DeltaNet paper ("Parallelizing Linear Transformers with the Delta Rule over Sequence Length", NeurIPS 2024) made delta-rule linear attention trainable at scale — the family now appears inside several production hybrid LLMs. Date marks the ecosystem's consolidation; DeltaNet paper June 2024.