OSWorld 2.0
evalYour notes
XLANG Lab's successor to OSWorld, the incumbent computer-use standard: 108 long-horizon real-world workflows (median ~1.6 hours of human time, ~318 tool calls per task vs ~30 in v1). The best frontier agent at release (Claude Opus 4.8) completes 20.6% vs 83.5% on OSWorld 1.0 — and the failures are constraint tracking, mid-task information processing, and hidden-state recovery, not UI control.
OSWorld 2.1 (September 16, 2026) pins tasks, assets, mocked websites and the VM image for reproducible runs, fixes evaluation bugs, and adds Gemini, Muse Spark and Qwen agents plus Claude Code, Codex, DeepSeek and Muse Code hybrid workflows. Labs now quote it in launch tables with partial-credit scoring: Anthropic reports Claude Opus 5.5 at 81.8% and Sonnet 5.5 at 80.1% (partial credit, not comparable to the 20.6% full-completion figure above).
Paper
Evaluation Details
Top Scores
| Model | Score | Date |
|---|---|---|
| Claude Opus 4.8 | 20.6% | 2026-06-28 |