Skip to content

What happened

Anthropic Fable benchmark rerun shows a sharp drop after reopening.

Why it matters

For builders who rely on model selection and benchmarking, the post highlights that a model’s ranking can change materially after a version change, especially when refusals replace answers on otherwise safe tasks. It also reinforces that benchmark results need to be checked against refusal rates and not just aggregate placement.

Severity rationale

[S3] benchmark-observed refusal regression on safe tasks

Source: abdullin.com
See also: Hallucination Incident Index