CL-bench: A Benchmark for Context Learning
evalYour notes
A benchmark for context learning, which its authors define as learning new knowledge from the context supplied at inference time and reasoning with it, as distinct from long-context tasks that mostly test retrieval and from in-context learning of simple task patterns from demonstrations. Domain experts built 500 contexts holding knowledge absent from pre-training, either newly created or drawn from niche, emerging long-tail material (for example the laws and case precedents of a fictional country, which the model must apply to adjudicate cases), and wrote 1,899 tasks on them with 31,607 verification rubrics, an average of 63.2 rubrics and about 20 hours of expert work per context. The tasks fall into four categories and 18 sub-categories: domain knowledge reasoning, rule system application, procedural task execution, and empirical discovery and simulation. Grading is strict: a task counts as solved only if the response passes every rubric, with GPT-5.1 as the judge; its verdicts agreed with Claude Opus 4.5 and Qwen3-Max verifiers, and with human checks of 100 samples, more than 90% of the time. Across ten frontier models in reasoning mode the paper found an average solving rate of 17.2%; the best, GPT-5.1 (high), reached 23.7%, followed by Claude Opus 4.5 at 21.1%, while Tencent's HY 2.0 scored 17.2% against o3's 17.8%. Empirical discovery was hardest, with solving rates around 11%.
Epoch AI added CL-bench to its Epoch Capabilities Index, describing it as created by the Tencent Hunyuan team and Fudan University's NLP group. On the project leaderboard at clbench.com, GPT-5.4 at xhigh effort leads at 27.9%, ahead of GPT-5.1 (high) at 23.7%; Tencent's Hy3 preview scores 22.8%. A sequel, CL-bench Life (arXiv 2604.27043, April 2026), applies the idea to messy real-life contexts such as group chats, personal notes and game logs, with 405 tasks and 5,348 rubrics (the best model solved 19.3% at release), and is also in the ECI. The 27-author paper lists the Hunyuan Team at Tencent and Fudan University; Shihan Dou, Ming Zhang and Zhangyue Yin are co-first authors, Pluto Zhou and Tao Gui are marked as corresponding authors, and Fudan's Xipeng Qiu and Xuanjing Huang are among the co-authors. The evaluation code is Apache 2.0, and the repository has about 580 GitHub stars.
Paper
Evaluation Details
Top Scores
| Model | Score | Date |
|---|---|---|
| GPT-5.4 (xhigh effort; clbench.com leaderboard) | 27.9% | — |
| GPT-5.1 (high; paper) | 23.7% | 2026-02-03 |
| Hy3 preview (clbench.com leaderboard) | 22.8% | — |