RecreationWorld and RecreationBench
evalYour notes
A five-platform environment framework and held-out benchmark for hybrid computer-use agents from Alibaba Token Hub, built on one task shape: recreation. The agent is given a running reference application and must discover its behavior by operating it, then build a faithful implementation, with no prescribed workflow — so GUI exploration, coding and visual verification interleave rather than stack. The reference doubles as an oracle: hidden behavioral tests derived from it give execution-grounded rewards. RecreationWorld supplies reproducible environments on Ubuntu, macOS, Windows, Android and Web plus a unified harness with native GUI control and coding tools, reading each platform through its own automation interface (AT-SPI, AXUIElement, UI Automation, UiAutomator, browser assertions). Code is MIT (89 stars at filing); the frozen task bundle ships as a HuggingFace dataset (5,222 downloads since it was posted on 2026-09-18).
RecreationBench is the evaluation half: 250 tasks, 50 per platform, each scored against two frozen inventories so a candidate cannot change its own denominator — Prog, assertions read through the platform's structured automation interface after replaying an interaction, and VLM, reference-grounded natural-language assertions on screenshots judged by a Qwen3.7-Plus judge at temperature zero. A candidate that cannot be built or launched scores zero. Every case is validated on the reference and by human reviewers before the suite is frozen. GPT-6 Astra leads at 58.06% overall (58.19 Prog / 57.92 VLM) but passes every programmatic test on just 2.80% of applications; Claude Opus 5 scores 44.16, GPT-5.6 Sol 42.06, Grok 4.6 36.73, Qwen3.8-Max-0902 34.80, Kimi K3 31.41, Claude Opus 4.8 31.10, GLM-5.3 24.38, Gemini 3.7 Flash 21.12 and Qwen3.7-Plus 9.15. Agents reproduce static interface structure far more reliably than interactions and computed outputs, and their applications stay smaller and more monolithic than the references. The paper also uses the framework as a data engine: Qwen3.8-Max generates recreation trajectories that rejection sampling against the behavioral verifiers trims to a balanced 35,000-trajectory SFT mixture (7,000 per platform), and both fine-tuned initializations (Qwen3.7-Plus and an in-house continued-pretrained Qwen-Flash checkpoint) end above their first measured checkpoint on five out-of-distribution benchmarks — ProgramBench, GameCraft-Bench, Vision2Web, OSWorld 2.0 and WeaveBench — while verifying their rendered output more often.
Paper
Evaluation Details
Top Scores
| Model | Score | Date |
|---|---|---|
| GPT-6 Astra | 58.06% | 2026-09-18 |
| Claude Opus 5 | 44.16% | 2026-09-18 |
| GPT-5.6 Sol | 42.06% | 2026-09-18 |