Skip to content

A wave of work argues that headline benchmark and leaderboard numbers systematically hide the variables that actually determine production behavior—harness design, quantization, sampling settings, and benchmark composition—amounting to an evaluation-validity crisis.

Evidence

  • Harness/scaffold is shown to be a dominant hidden variable: the Scaffold Effect reports up to 40x token differences by harness choice, and ‘API Settings and Harness Design’ plus Boris Cherny’s ‘delete your harness every six months’ reinforce that scores measure systems, not models.
  • Aggregate scores mask sample- and channel-level failure: quantized agents’ error budget flattens tool-calling degradation, and ‘Benchmarks Are Not Monolithic’ shows aggregate accuracy hides what individual samples test.
  • Validity does not transfer: the ‘non-composition principle’ warns benchmark inference chains break when assumptions/data/populations shift, and WorkSurface-Bench isolates routing failures that answer-accuracy metrics conflate.
  • Layer- and source-specific RAG evaluation (LayerRAG-Bench, knowledge-inconsistency detection) shows surface groundedness checks miss stale evidence, permission, and session-state errors.

Implications

  • Model comparisons and deployment go/no-go decisions require controlling for harness, quantization, and API settings, plus sample-level and layer-level auditing rather than single aggregate scores.
  • Expect demand for reproducible reference harnesses (e.g., SimpleWikiSearch) and per-dimension diagnostics as standard practice in eval pipelines.

Concepts

LLM Evals RAG Evaluation Code Agents Tool Use Agents

Confidence

high