NeurIPS 2025 Oral: query-agnostic KV-cache compression — scores each KV pair by its usefulness for reconstructing the context rather than for answering a specific query, so one compressed cache serves arbitrary future queries and cache reuse. Achieves 3–4x cache reduction with negligible quality loss across QA, retrieval, reasoning, and code tasks. From Hyun Oh Song's MLLab (snu-mllab).

Paper

Venue NeurIPS 2025 (Oral)
efficiencyinference

Related