Thunder-DeID
dataset Your tags
Your notes
EMNLP 2025 Findings: the first Korean legal de-identification dataset — ~27K annotated entities across 595 fine-grained types over court judgments — paired with SNU_Thunder-DeID encoder models (340M/750M/1.5B) that set state of the art on court-judgment de-identification. Part of the Thunder group's Korean data stack around Thunder-LLM.
Paper
Dataset
Size ~27K entities, 595 types