On-device counterpart to the Kanana-2-30B MoE line: a 3B trained from scratch on TPUs and a 1.3B built via cascade pruning + distillation from it, with 3:1 sliding-window hybrid attention (~72.7% KV-cache cut at 32K) and a new tokenizer (+30% Korean efficiency). Kanana Open License (commercial use OK); a 0.9B is mentioned but unreleased.

Model Details

Architecture DENSE

Variants

Name Parameters Notes
kanana-2-3b 3B From scratch on TPUs
kanana-2-1.3b 1.3B Cascade pruning + distillation from the 3B
on-deviceopen-weight

Related