Tao Yu's XLANG lab moves from computer-use agents into robotics: "FineVLA: Fine-Grained Instruction Alignment for Steerable VLA Policies", with a vision-language-action model built on a Qwen3.5-VL 397B-A17B MoE backbone (~17B active) plus the RoboFine-bench companion. The lab's first VLA line, extending the OSWorld/OpenCUA agent stack to physical manipulation.

Paper

roboticsmultimodalopen-weight