Enterprise text-to-SQL workflow benchmark (ICLR 2025 Oral) from the XLANG Lab: 632 real-world problems derived from enterprise database use cases — databases with 1,000+ columns on BigQuery/Snowflake, multiple SQL dialects, and solutions often exceeding 100 lines across metadata search, dialect docs, and project codebases. At release, an o1-preview-based code agent solved only 21.3% of tasks, versus 91.2% on predecessor Spider 1.0 and 73.0% on BIRD — making it a standard eval for agentic coding and data-engineering ability.

Successor to Spider (2018), the field's canonical academic text-to-SQL benchmark, both led by Tao Yu.

Paper

Venue ICLR 2025

Evaluation Details

Tasks 632
Scoring execution-based
benchmarkevaluationcodingagentic

Related