CyberGym-E2E
evalYour notes
The successor to CyberGym from Berkeley's Center for Responsible, Decentralized Intelligence (RDI), led by Dawn Song, which moves from reproducing known vulnerabilities to the whole defensive loop. Each of its 920 real-world vulnerabilities across 139 open-source projects, drawn from Google's OSS-Fuzz, places an agent inside the project's own build environment with the vulnerable codebase, build scripts and test scripts. In the end-to-end setting the agent gets no ground-truth data: it must find a bug, write a proof-of-concept input that triggers a sanitizer crash, and patch the code. Four cumulative checks grade the result: the PoC crashes the unpatched build (S1), the patch removes that crash (S2), the project's developer-written functionality tests still pass (S3, the success criterion), and, as a diagnostic, the patch also fixes the specific ground-truth vulnerability (S4). A patch-only setting instead hands the agent the ground-truth PoC and crash log to isolate repair. Tasks come from an automated, agent-assisted pipeline that finds each fix commit, rebuilds the vulnerable and patched versions on modern toolchains and has a coding agent set up each project's unit tests, with a human expert validating test coverage at the end, so the set can keep ingesting new OSS-Fuzz vulnerabilities.
Discovery is the bottleneck. On the initial 615-task set, with a $10 and 90-minute budget per task, Claude Opus 4.5 in Claude Code fixed 82.3% of tasks when given the PoC but succeeded end-to-end (S3) on only 19.2%, and no configuration exceeded 22.6% (Gemini 3 Pro in Gemini CLI). On the full 920 tasks GPT-5.4 in Codex reached 65.9% end-to-end and Claude Opus 4.6 62.6% with the cost cap lifted, while S4 topped out at 26.2%, because agents often patch a real vulnerability other than the ground-truth one. On 28 September 2026 Artificial Analysis made a 131-task subset (one task per project, selected for difficulty and filtered for oracle leaks and sandbox compatibility), CyberGym-E2E-AA, one of three equally weighted components of its new Cyber Index, alongside Collinear AI's CWE-Bench-AA and Vercel's DeepsecBench-AA. There MiMo-V2.6-Pro leads at 78.6% pass@1, ahead of GPT-6 Luna at max effort (77.9%) and Grok 4.7 at xhigh (74.0%), while GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B and Qwen3.8 27B refuse at least 98% of tasks. The paper (16 authors, 12 of them at UC Berkeley, with others from Johns Hopkins, UC Santa Cruz and UC Santa Barbara; co-first authors Tianneng Shi and Robin Rheem) is an ICML 2026 paper. Apache 2.0; about 80 GitHub stars, against about 920 for the original CyberGym.
Paper
Evaluation Details
Top Scores
| Model | Score | Date |
|---|---|---|
| MiMo-V2.6-Pro (CyberGym-E2E-AA, 131 tasks, pass@1) | 78.6% | — |
| GPT-6 Luna (max effort; CyberGym-E2E-AA) | 77.9% | — |
| GPT-5.4 (Codex; paper, 920 tasks, S3) | 65.9% | — |
| Claude Opus 4.6 (Claude Code, no cost cap; paper, 920 tasks, S3) | 62.6% | — |