ITBench
evalYour notes
IBM Research's benchmark for AI agents on enterprise IT automation, covering three personas: Site Reliability Engineering (diagnosing and resolving incidents in Kubernetes environments), Compliance and Security Operations (CISO: assessing compliance posture against CIS-benchmark controls) and Financial Operations (cost overruns and anomalies). The first release held 94 scenarios (42 SRE, 50 CISO, 2 FinOps); the SRE scenarios are modeled on real incidents from IBM's own SaaS products, and runs are scored on pass@1 plus domain metrics such as topology-aware fault-localization and propagation-chain scores and mean time to diagnosis and repair. Agents powered by the state-of-the-art models of the time (GPT-4o, Llama and Granite baselines) resolved only 13.8% of SRE scenarios, 25.2% of CISO scenarios and none of the FinOps ones. At release IBM open-sourced 11 of the 94 scenarios with its baseline SRE and CISO agents and kept the rest to evaluate submitted agents through a managed leaderboard. The paper (43 authors, nearly all from IBM Research with five from the University of Illinois Urbana-Champaign; co-first authors Saurabh Jha, Rohan Arora and Yuji Watanabe) was an oral at ICML 2025. Apache 2.0, about 500 GitHub stars.
Adoption came through leaderboards. IBM added ITBench to Kaggle's enterprise AI leaderboards in December 2025, and on 27 May 2026 Artificial Analysis and IBM Research launched ITBench-AA, AA's own implementation of the SRE track: 59 Kubernetes incident root-cause tasks (40 from IBM's public release and 19 private tasks from the ITBench team), each an offline snapshot of alerts, events, traces, metrics, logs and application topology that the agent inspects with a single shell tool in AA's Stirrup harness (100-turn cap, three repeats per task). A repeat scores zero if it misses any ground-truth root-cause entity and otherwise earns the precision of its submitted entities. Every model scored below 50% at launch; GPT-5.6 Sol at max effort now leads at 56.2%, ahead of StepFun's Step 5 Preview at 55.6%. ITBench-AA is a standalone AA evaluation rather than part of the Intelligence Index, and ITBench is not among the benchmarks in Epoch's Capabilities Index.
Paper
Evaluation Details
Top Scores
| Model | Score | Date |
|---|---|---|
| GPT-5.6 Sol (max effort; ITBench-AA) | 56.2% | — |
| Step 5 Preview (ITBench-AA) | 55.6% | — |
| GPT-4o (baseline SRE-Agent, SRE diagnosis pass@1 at release) | 13.8% | 2025-02-07 |