LLM-jp-4
modelYour notes
Latest generation with MoE and "thinking" variants. 32B/3.8B active MoE (128 routed experts, top-8, 32 layers, 2560 hidden, 40 heads). 65K context. Trained on 11.7T tokens with llm-jp-tokenizer v4.0. Apache 2.0.
Also includes dense 8B (9B params) with base and thinking variants. MT-Bench JA: 7.57-7.82, MT-Bench EN: 7.70-7.86. Evaluated using GPT-5.4 as judge. Built by the NII LLM Research Center, which UCL's Pontus Stenetorp co-leads — joint NII/UCL attribution. A dense 33B (64 layers, hidden size 5,120) joined the series in August 2026, trained on the same 11.7T-token pre- and mid-training pipeline.
LLM-jp-4.1 (thinking models for all three sizes on HuggingFace September 15–16, 2026; technical blog September 28) redoes only the post-training. The SFT set grows to about 12B tokens, trained for three epochs (~35B tokens, about three times LLM-jp-4's), with far more STEM data whose reasoning and answers were regenerated by gpt-oss-120b, and adds the series' first tool calling (NVIDIA Nemotron agentic data plus a Japanese translation); DPO follows, with no RL. Average scores rise over LLM-jp-4 at every size, mostly in math, science and instruction following, and safety (AnswerCarefully) improves while MT-Bench stays flat or dips as answers get shorter. On NII's averages the 33B thinking model beats Olmo-3.1-32B-Think and Qwen3-32B and roughly matches gpt-oss-120b, but trails Qwen3.8-27B, Gemma-4-31B-it and Muse-Glimmer-30B. The SFT and DPO data are released alongside the Apache 2.0 weights; the 33B SFT run took about 67 hours on 16 nodes of 8×H200.
Model Details
Variants
| Name | Parameters | Notes |
|---|---|---|
| llm-jp-4-8B | 8B | — |
| llm-jp-4-32B-A3B | 32B | — |
| llm-jp-4-33B | 33B | Dense, 64 layers, hidden 5,120, 40 heads, 65K context; base and thinking models released August 14, 2026 |
| llm-jp-4.1 (8B / 32B-A3B / 33B thinking) | — | September 15-16, 2026: post-training refresh of the same bases (larger STEM-heavy SFT, first tool calling, DPO, no RL); SFT and DPO data released |