AREX-2
modelYour notes
The second AREX model from BAAI's AREX Team (project leader Zheng Liu, code under the VectorSpaceLab GitHub org), aimed at test-time self-improvement: given more rounds on a task, the agent should turn them into a better solution by its own judgment. The paper splits this into reflection, which sets how much a round gains, and long-horizon execution, which sets how many rounds stay productive, and argues that both are domain-agnostic skills best learned where feedback is unambiguous. A teacher model turns GitHub machine-learning repositories and online-judge problems into scored environments, kept only if the reference solution runs and a simple baseline scores well below it. A teacher agent then works each task for hours and hundreds of tool calls, reading code and papers and running experiments between scored submissions. Trajectories are kept or dropped as a whole by final score, so failed rounds and regressions stay in, and the loss falls on the steps that recover and make progress. AREX-2 is a supervised fine-tune of Qwen3.8-27B on these trajectories plus the unchanged deep-research data of the first AREX recipe: a 27B dense multimodal model with a 262,144-token context, under Apache 2.0.
It scores 81.8 on MLE-bench Lite (Any Medal, mean of three seeds, run with skills in context), 9.1 points above GPT-5.6 Sol, and 70.7 on the 188-task Frontier-CS Agent Track, 16.0 points above the best open-weight baseline and 5.7 below GPT-5.6 Sol. With no new search data it beats both earlier AREX models in deep research: BrowseComp 84.0, text-only HLE 52.6, GAIA 92.2 and DeepSearchQA 93.8, against 82.5, 52.4, 85.4 and 89.9 for the 122B AREX-Base. The gains come from sustained iteration. On Frontier-CS it reaches 54.4 after one hour, 65.9 after two and 70.7 after five, while DeepSeek-V4-Pro stops at 44.7 after two hours; on BrowseComp, with no correctness feedback, accuracy climbs from 64.8 at 47 turns to 84.0 at 143. In a stage-wise ablation on MLE-bench Lite, skills and more rounds lift the untrained base from 28.8 to 68.2, and training adds 13.6 points on top. BAAI released the weights and the paper PDF on Hugging Face on September 29, 2026 (no arXiv ID at filing); the GitHub repository holds the evaluation runners and dataset definitions for the six benchmarks.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| MLE-bench Lite | 81.8 | Any Medal, mean of 3 seeds, with skills |
| Frontier-CS (Agent Track, 188 tasks) | 70.7 | — |
| BrowseComp | 84.0 | — |
| HLE (text-only) | 52.6 | — |
| GAIA | 92.2 | — |
| DeepSearchQA | 93.8 | F1 |
Variants
| Name | Parameters | Notes |
|---|---|---|
| AREX-2 | — | 27B dense; supervised fine-tune of Qwen3.8-27B |