VC-Tooler: Learning Compositional and Adaptive Visual Tool Use
paper Your tags
Your notes
Trains VLMs to use visual tools as a compositional and adaptive capability rather than fixed-tool-space grounding: a hierarchical trajectory-synthesis pipeline builds a bank across three levels (single-tool grounding, multi-tool composition, diverse tool contexts/interfaces), then two-stage training — supervised cold start followed by RL rewarding accurate, efficient, context-aware tool calls. State-of-the-art among open models on general and agentic benchmarks (95.8% on V*, 35.3% on VTC-Bench), with transfer to richer tool settings at inference. SFT and RL trajectory datasets released on HuggingFace and ModelScope. Alibaba Group (affiliation per the project page; senior author Jieping Ye).