What happened
Anthropic Fable benchmark rerun shows a sharp drop after reopening.
Why it matters
For builders who rely on model selection and benchmarking, the post highlights that a model’s ranking can change materially after a version change, especially when refusals replace answers on otherwise safe tasks. It also reinforces that benchmark results need to be checked against refusal rates and not just aggregate placement.
Severity rationale
[S3] benchmark-observed refusal regression on safe tasks
Source: abdullin.com
See also: Hallucination Incident Index