Real-time vision-language understanding of effectively infinite video streams: attention sinks plus a short-vision/long-text KV window give ~8 FPS on a single H100 while beating GPT-4o-mini on the accompanying long-video benchmark — the StreamingLLM idea carried into multimodality.

Paper

multimodalefficiencyopen-weight