The de-facto industry benchmark for computer-use agents (NeurIPS 2024), from Tao Yu's XLANG Lab: a scalable real-computer environment (Ubuntu/Windows/macOS) with 369 execution-verified tasks spanning real web and desktop apps. At launch humans solved 72.4% vs. 12.2% for the best agent; it became the metric quoted in Anthropic and OpenAI computer-use announcements, later hardened as OSWorld-Verified.

OSWorld 2.0 (June 2026, arXiv 2606.29537) raises the ceiling to 108 long-horizon real-world workflows averaging ~318 tool calls per task. Companions from the same lab: OSWorld-G grounding benchmark + CUA-Gym RLVR training pipeline (NeurIPS 2025 Spotlight). HKU-led multi-institution effort.

Paper

Venue NeurIPS 2024

Evaluation Details

Tasks 369
Scoring execution-based
benchmarkevaluationagentic

Related