An Apple approach to long context in which a vision-language model reads text as rendered page images instead of tokens. Because an image encoder maps a fixed-size image to a fixed number of visual tokens, the rendering resolution works as a compression knob, but accuracy collapses once the characters shrink below what the encoder can resolve. LensVLM (an inference framework plus a post-training recipe) lets the model scan the compressed images and then expand only the relevant pages back to their uncompressed form through learned tools, much as an agent pages content back into its context.

Built on Qwen3.5-9B-Base, it stays comparable to the full-text upper bound at 4.3× effective compression and beats retrieval-based, text-compression and visual-compression baselines up to 10.1× effective compression across seven text QA benchmarks; it also carries over to multimodal document and code understanding, with the margin over baselines growing as compression rises. The analysis finds that training makes visual compression robust to rendering choices, that the model leans more on expanded content as compression grows, and that expanding to text suits rendered text while high-resolution image expansion suits native documents whose layout matters. The paper (first author Roy Xie, with Bhuwan Dhingra; both Apple and Duke University) went to arXiv on May 7, 2026; the date above marks the September 21 release of the LensVLM-9B weights on HuggingFace under the Apple Machine Learning Research Model License, with code under the Apple Sample Code License.

Model Details

License Apple Machine Learning Research Model License
Base model qwen3.5

Variants

Name Parameters Notes
LensVLM-9B — 9B VLM post-trained from Qwen3.5-9B-Base; supports 5x, 10x and 15x text-compression settings in the demo

Paper

Authors: Roy Xie · Dan Friedman · Donghan Yu · Bowen Pan · Christopher Fifty · Jang-Hyun Kim · Xianzhi Du · Zhe Gan
multimodalvisionlong-contextefficiencyopen-weightresearch

Related