Fine-grained Language model evaluation based on Alignment SKill sets (ICLR 2024): decomposes instruction-following evaluation into 12 instance-specific skills instead of a single holistic score, showing skill-level scoring improves both human–model and model–model evaluator agreement. The methodological start of the LK Lab evaluator line that produced Prometheus and BiGGen Bench.

Paper

Venue ICLR 2024
evaluation

Related