Skip to content

Outcome-only and headline-benchmark evaluation is being exposed as systematically blind, driving a shift toward process-, resource-, and provenance-level evaluation.

Evidence

  • Trajectory-Judge shows outcome-only evaluation misses 55% of silent agent faults (correct output from incorrect logic)
  • Interface-Induced Trajectory Censoring shows well-formed tool calls silently dropped by parsers, making capability and misconfiguration indistinguishable in scores
  • Comparing Agentic Systems and READY argue similar completion rates hide order-of-magnitude differences in resource cost and oversight requirements
  • BenchMIRT finds safety benchmarks like BBQ may measure general reasoning, and ‘Logarithmic scales hide the cost chasm’ shows visualizations mask true cost-performance gaps

Implications

  • Deployment decisions require step-level rubrics, execution provenance, and resource instrumentation, not leaderboard accuracy alone
  • Reported agent capability numbers should be treated as fragile until interface, cost, and process faults are ruled out

Concepts

LLM Evals Agents Tool Use Code Agents

Confidence

high