AWQ
paper Your tags
Your notes
Activation-aware Weight Quantization (MLSys 2024 Best Paper, ~3.6K★): protect the ~1% of salient weights identified by activation statistics and 4-bit quantize the rest. Still the default 4-bit path in vLLM, HF Transformers, and TensorRT-LLM through 2026 — the canonical artifact of the HAN Lab quantization line (preprint June 2023; dated by its award venue). Successors QServe (W4A8KV4) and LServe carry the line into serving-system co-design (MLSys 2025).