Covering Jul 06, 2026 to Jul 12, 2026 (UTC) · generated Jul 13, 2026.
This week’s radar flagged 3 hallucination-adjacent incidents across 9 named models (-1 vs. the previous 4-incident week). Severity split: 0 S1, 0 S2, 2 S3, 1 S4. By category: 1 hallucination, 0 jailbreak, 2 refusal, 0 bias. Verification funnel: 7 flagged by the heuristic → 3 verified by the LLM judge → 4 rejected as non-incidents (announcements, tutorials).
Severity breakdown
| Severity | Count | % of week |
|---|---|---|
| S1 | 0 | 0% |
| S2 | 0 | 0% |
| S3 | 2 | 67% |
| S4 | 1 | 33% |
Category breakdown
| Category | Count | % of week |
|---|---|---|
| Hallucination | 1 | 33% |
| Jailbreak | 0 | 0% |
| Refusal | 2 | 67% |
| Bias | 0 | 0% |
Top models by incidents
9 distinct model names were mentioned across this week’s incidents; the top 9 by incident count are ranked below.
| # | Model | Incidents | Δ vs. previous week |
|---|---|---|---|
| 1 | ChatGPT | 1 | +1 |
| 2 | Claude | 1 | +1 |
| 3 | Fable 5 | 1 | -2 |
| 4 | GPT-5.0 | 1 | 0 |
| 5 | GPT-5.1 | 1 | 0 |
| 6 | Gemini | 1 | +1 |
| 7 | Mythos | 1 | +1 |
| 8 | Opus 4.8 | 1 | -1 |
| 9 | Sonnet 5 | 1 | +1 |
Most severe incidents this week
- [S3] ✓ Verified — Opinionated notes on AI-assisted testing and agentic coding — hallucination · GPT-5.0, GPT-5.1 · confidence: medium · global importance 3/5 (Jul 08, 2026). For builders, the useful signal is not that agents are magical, but that they can speed up testing and bug investigation while also producing plausible but false outputs. That makes verification discipline central when using coding agents in real workflows. danluu.com
- [S3] ✓ Verified — Rob Patro says Fable was not useful for a Rust rewrite task — refusal · Fable 5, Mythos, Opus 4.8 · confidence: medium · global importance 2/5 (Jul 08, 2026). The post is a concrete report of an AI model refusing a code rewrite task because of safety classification, then declining to explain how to rephrase the prompt. For builder workflows, it is a reminder that model gating can block normal software work even when the user is asking for a straightforward code port. combine-lab.github.io
- [S4] ✓ Verified — A Claude user says recent behavior feels more pushy and less consistent — refusal · Claude, Gemini, ChatGPT, Sonnet 5 · confidence: medium · global importance 2/5 (Jul 11, 2026). This is a user-level report about perceived changes in Claude’s conversational behavior, with direct comparison against Gemini and ChatGPT. For builders, it is mainly a signal that assistant behavior consistency and refusal handling can materially affect whether users trust a tool in long-form workflows. androidauthority.com
Methodology
This is an MVP proxy index, not a verified incident registry. There is no dedicated hallucination-incident case database in GROUNDING yet — an “incident” here is any record from the AI-news radar’s own daily analysis journal (data/results/, ~2,500 records/week of AI news, papers, and community posts) that (1) names at least one model, (2) matches a hallucination / jailbreak / refusal / bias keyword pattern in its title, summary, or topics, and (3) was filed under a field-report category (model release, industry, or opinion coverage) rather than an academic research paper proposing a detection/mitigation method — the latter are excluded on purpose so this stays a field-incident signal, not a synthetic-benchmark leaderboard.
Consequences worth knowing before citing a number from this page: volume is low by construction (typically ~10-15 qualifying incidents/week); model names are raw NER extractions passed through a curated canonicalizer (known vendor-prefix/formatting duplicates are merged, e.g. “Claude Opus 4.8” and “Opus 4.8”), but genuinely ambiguous bare mentions spanning several concurrently-discussed versions (e.g. “Opus”) are left as their own entry rather than guessed onto one version — week-over-week deltas are still computed on the canonical string only; and “case links” point at the original source article, since no dedicated per-incident page exists yet. Severity S1-S4 is derived from the item’s global_importance score (1-5): 5→S1, 4→S2, 3→S3, 1-2→S4. Confidence (high/medium/low) reflects whether the keyword matched in the title, the prose, or only the tags.
Verification layer. Every candidate above is additionally run through an LLM judge (src.report.incident_verify, not a human reviewer) that decides whether it is a real field incident versus a research paper, announcement, or tutorial, and rechecks its category, severity, and named models. Candidates the judge rejects are removed from this report entirely; candidates it confirms are marked ”✓ Verified”, shown with the judge’s rechecked severity and category (which override the heuristic’s guess in every count above), and get a full write-up as a Case page under “/incidents/Cases/“. Everything else — not yet judged, or judged “uncertain” — is marked “Candidate”: still a keyword match, not yet independently verified by either the judge or a person.
Full weekly snapshot: Week of 2026-07-06 archive.