"The Open-Ended Homogeneity of Language Models" — NeurIPS 2025 Best Paper (Datasets & Benchmarks). Introduces the Infinity-Chat dataset and shows models collapse to homogeneous outputs on open-ended prompts, even across different model families — a diversity/mode-collapse problem directly relevant to eval design and synthetic-data pipelines. UW-led (Jiang, Tsvetkov) with CMU, Stanford, and AI2.

A September 2026 reevaluation from Stanford (Schaeffer, Miranda, Kazdan, Chudnovsky, Koyejo; arXiv 2609.33936) disputes the evidence. The flagship "metaphor involving time" responses show one dominant vehicle plus a long tail rather than two clusters; under a stricter null (same-prompt answers expressing different ideas) 20–32% of pairs already exceed the paper's 0.8 convergence threshold, though same-prompt pairs still cross it two to three times as often; and simple prompting reliably raises measured diversity. The authors conclude the published evidence does not establish the effect, without ruling it out.

Paper

evalresearch