SNU Korean Eval Suite
eval Your tags
Your notes
Localized Korean benchmark set from the Thunder Research Group: SNU_Ko-GSM8K, Ko-IFEval, Ko-ARC, Ko-EQ-Bench, Ko-WinoGrande, Ko-LAMBADA, and Ko-MuSR — careful Korean localizations of standard reasoning/instruction benchmarks rather than raw translations — plus the Thunder-NUBench / Thunder-KoNUBench negation-understanding benchmarks (the Korean one evaluated across 47 LLMs). Built to evaluate Thunder-LLM and other Korean models, released 2024-12 through 2025.