KVzip
paper Your tags
Your notes
NeurIPS 2025 Oral: query-agnostic KV-cache compression — scores each KV pair by its usefulness for reconstructing the context rather than for answering a specific query, so one compressed cache serves arbitrary future queries and cache reuse. Achieves 3–4x cache reduction with negligible quality loss across QA, retrieval, reasoning, and code tasks. From Hyun Oh Song's MLLab (snu-mllab).