"Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model" — a 4B VLM that allocates visual tokens like codec bits: >75% fewer visual tokens and up to 3.5× speedup while beating Qwen3-VL-4B on all reported video benchmarks, with an event-gated live-commentary mode. The Mage-ViT encoder is trained from scratch; the language decoder is Qwen3-4B. Apache 2.0; 435K downloads in its first ten days.

Model Details

Architecture DENSE
Base model qwen3

Paper

multimodalvisionopen-weightefficiency