BRIGHT
eval Your tags
Your notes
The first reasoning-intensive retrieval benchmark (ICLR 2025): 1,384 real-world queries across diverse domains (economics, psychology, mathematics, coding) where relevant documents cannot be found by surface-form matching — the top MTEB embedder at the time (SFR-Embedding-Mistral, 59.0 nDCG@10 on MTEB) dropped to 18.0 nDCG@10. Now a standard hard-set in embedding and retriever evaluations, and the reference eval for the "reasoning retriever" model class.
XLANG Lab + multi-institution collaboration, HKU-led. Companion to hkunlp's legacy Instructor embedding line.
Paper
Evaluation Details
Questions 1,384
Scoring nDCG@10