GuidedQuant
paper Your tags
Your notes
ICML 2025: post-training LLM quantization that injects end-loss gradient information into the layer-wise quantization objective, improving weight-only scalar, weight-only vector, and weight-and-activation quantization at 2–4 bits. From Hyun Oh Song's MLLab, with the Q-Palette follow-up (fractional-bit quantizers, NeurIPS 2025) extending the line.