Polaris
model Your tags
Your notes
An open RL post-training recipe for advanced reasoning from Lingpeng Kong's hkunlp (lead: Chenxin An), built around calibrated data difficulty, staged temperature/length scheduling, and multi-stage RL on Qwen3 bases. The Polaris-4B-Preview headline result — a 4B model rivaling much larger commercial reasoners on AIME-class math — made it a reference recipe for scaling RL on small open models.
Polaris V2 (March 2026) scales the recipe to Qwen3-235B-A22B alongside refreshed Qwen3-4B checkpoints. Released as blog + checkpoints; no arXiv paper yet. Joint work with ByteDance Seed and Fudan collaborators, HKU-led.
Model Details
Base model qwen3
Variants
| Name | Parameters | Notes |
|---|---|---|
| Polaris-4B-Preview | — | RL post-trained from Qwen3-4B |
| Polaris-7B-Preview | — | RL post-trained from a 7B base |
| Polaris-V2-Qwen3-235B-A22B | — | V2 recipe on Qwen3-235B-A22B (2026-03) |