Qwen3.8-2.4T-A95B
model Your tags
Your notes
Alibaba's first open-weight Qwen-Max-class release and its largest public model: a 2.4T-parameter MoE with 95B active per token. Its 92-layer hybrid stack interleaves three Gated DeltaNet linear-attention blocks with one gated-attention block; 10 of 512 routed experts plus one shared expert are active. Native context is 262K and can extend to 1.01M tokens.
Built for coding, professional work, research, and long-horizon agents, with controllable reasoning effort and preserved multi-turn thinking. The weights use the custom Qwen3.8-Max license: modification and redistribution are allowed, but large consumer products must display attribution and qualifying MaaS/work- assistant businesses must obtain a separate commercial license.
Model Details
Architecture MOE
Parameters 2.4T
Active params 95B
Experts 512 (top-10)
Context window 1,010,000
AA Intelligence 58
License Qwen3.8-Max
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| Terminal-Bench 2.1 | 86.6% | — |
| SWE-bench Pro | 67.7% | — |
| GPQA Diamond | 92.6% | — |