A benchmark for context learning, which its authors define as learning new knowledge from the context supplied at inference time and reasoning with it, as distinct from long-context tasks that mostly test retrieval and from in-context learning of simple task patterns from demonstrations. Domain experts built 500 contexts holding knowledge absent from pre-training, either newly created or drawn from niche, emerging long-tail material (for example the laws and case precedents of a fictional country, which the model must apply to adjudicate cases), and wrote 1,899 tasks on them with 31,607 verification rubrics, an average of 63.2 rubrics and about 20 hours of expert work per context. The tasks fall into four categories and 18 sub-categories: domain knowledge reasoning, rule system application, procedural task execution, and empirical discovery and simulation. Grading is strict: a task counts as solved only if the response passes every rubric, with GPT-5.1 as the judge; its verdicts agreed with Claude Opus 4.5 and Qwen3-Max verifiers, and with human checks of 100 samples, more than 90% of the time. Across ten frontier models in reasoning mode the paper found an average solving rate of 17.2%; the best, GPT-5.1 (high), reached 23.7%, followed by Claude Opus 4.5 at 21.1%, while Tencent's HY 2.0 scored 17.2% against o3's 17.8%. Empirical discovery was hardest, with solving rates around 11%.

Epoch AI added CL-bench to its Epoch Capabilities Index, describing it as created by the Tencent Hunyuan team and Fudan University's NLP group. On the project leaderboard at clbench.com, GPT-5.4 at xhigh effort leads at 27.9%, ahead of GPT-5.1 (high) at 23.7%; Tencent's Hy3 preview scores 22.8%. A sequel, CL-bench Life (arXiv 2604.27043, April 2026), applies the idea to messy real-life contexts such as group chats, personal notes and game logs, with 405 tasks and 5,348 rubrics (the best model solved 19.3% at release), and is also in the ECI. The 27-author paper lists the Hunyuan Team at Tencent and Fudan University; Shihan Dou, Ming Zhang and Zhangyue Yin are co-first authors, Pluto Zhou and Tao Gui are marked as corresponding authors, and Fudan's Xipeng Qiu and Xuanjing Huang are among the co-authors. The evaluation code is Apache 2.0, and the repository has about 580 GitHub stars.

Paper

Authors: Shihan Dou · Ming Zhang · Zhangyue Yin · Chenhao Huang · Yujiong Shen · Junzhe Wang · Pluto Zhou · Tao Gui

Evaluation Details

Tasks 1,899
Domains 4
Scoring Rubric-based and binary per task: a response counts as solved only if it passes every one of the task's expert-written rubrics, judged by GPT-5.1; the headline metric is the task solving rate
Saturation Not saturated: best 27.9% (GPT-5.4 at xhigh effort, project leaderboard); 23.7% best at release
Used in: Epoch Capabilities Index
Domains: domain knowledge reasoning, rule system application, procedural task execution, empirical discovery and simulation

Top Scores

Model Score Date
GPT-5.4 (xhigh effort; clbench.com leaderboard) 27.9% —
GPT-5.1 (high; paper) 23.7% 2026-02-03
Hy3 preview (clbench.com leaderboard) 22.8% —
benchmarkevaluationlong-contextreasoning

Related