Outcome-only and headline-benchmark evaluation is being exposed as systematically blind, driving a shift toward process-, resource-, and provenance-level evaluation.
Evidence
- Trajectory-Judge shows outcome-only evaluation misses 55% of silent agent faults (correct output from incorrect logic)
- Interface-Induced Trajectory Censoring shows well-formed tool calls silently dropped by parsers, making capability and misconfiguration indistinguishable in scores
- Comparing Agentic Systems and READY argue similar completion rates hide order-of-magnitude differences in resource cost and oversight requirements
- BenchMIRT finds safety benchmarks like BBQ may measure general reasoning, and ‘Logarithmic scales hide the cost chasm’ shows visualizations mask true cost-performance gaps
Implications
- Deployment decisions require step-level rubrics, execution provenance, and resource instrumentation, not leaderboard accuracy alone
- Reported agent capability numbers should be treated as fragile until interface, cost, and process faults are ruled out
Concepts
LLM Evals Agents Tool Use Code Agents
Confidence
high