The load-bearing negative result for the self-evolving harness trend: current harness-evolution evaluations conflate genuine design gains with extra search compute. At matched compute budgets, automatic harness evolution does not consistently beat simple test-time scaling, and generalization to held-out tasks is weak. By Wang, Hajishirzi, Tsvetkov, Dasigi et al. (AI2 + University of Washington). The skeptical anchor for a subfield that produced three papers in July 2026 alone.

Paper

agentsevalresearch