W4A8KV4 quantization co-designed with the serving system (MLSys 2025, ~849★); LServe adds unified sparse attention for long context. The systems continuation of the AWQ line, alongside the lab's long-context attention work (DuoAttention ICLR 2025, XAttention ICML 2025, Radial Attention NeurIPS 2025).

Paper

Library

efficiencyinfrastructure

Related