An open RL post-training recipe for advanced reasoning from Lingpeng Kong's hkunlp (lead: Chenxin An), built around calibrated data difficulty, staged temperature/length scheduling, and multi-stage RL on Qwen3 bases. The Polaris-4B-Preview headline result — a 4B model rivaling much larger commercial reasoners on AIME-class math — made it a reference recipe for scaling RL on small open models.

Polaris V2 (March 2026) scales the recipe to Qwen3-235B-A22B alongside refreshed Qwen3-4B checkpoints. Released as blog + checkpoints; no arXiv paper yet. Joint work with ByteDance Seed and Fudan collaborators, HKU-led.

Model Details

Base model qwen3

Variants

Name Parameters Notes
Polaris-4B-Preview RL post-trained from Qwen3-4B
Polaris-7B-Preview RL post-trained from a 7B base
Polaris-V2-Qwen3-235B-A22B V2 recipe on Qwen3-235B-A22B (2026-03)
reasoningpost-trainingopen-weight