BiGGen Bench
eval Your tags
Your notes
Principled LM-as-judge benchmark: 77 tasks with instance-specific evaluation criteria rather than coarse task-level rubrics, spanning nine capabilities. Won a NAACL 2025 Best Paper Award (one of two among ~2,000 submissions, confirmed on the official awards page) — the capstone of the FLASK / Prometheus evaluator line. Joint work between the LK Lab (Minjoon Seo) and LG AI Research.
Paper
Evaluation Details
Tasks 77