69B-A3B MoE replacing dense MLA with LongCat Sparse Attention (LSA) — extending DeepSeek's DSA with streaming-aware, cross-layer, hierarchical indexing — at native 1M context. SWE-bench Verified 68.2 vs 54.4 for the dense sibling. The paper (Aug 3) doubles as the attention technical report for LongCat-2.0.

Model Details

Architecture MOE
Parameters 69B
Active params 3B
Context window 1,000,000

Paper

moeattentionopen-weightefficiency

Related