StreamingVLM
model Your tags
Your notes
Real-time vision-language understanding of effectively infinite video streams: attention sinks plus a short-vision/long-text KV window give ~8 FPS on a single H100 while beating GPT-4o-mini on the accompanying long-video benchmark — the StreamingLLM idea carried into multimodality.