EMNLP 2025 Findings: the first Korean legal de-identification dataset — ~27K annotated entities across 595 fine-grained types over court judgments — paired with SNU_Thunder-DeID encoder models (340M/750M/1.5B) that set state of the art on court-judgment de-identification. Part of the Thunder group's Korean data stack around Thunder-LLM.

Paper

Dataset

Size ~27K entities, 595 types
datamultilingual

Related