Data note
Fewer than 3 incidents matched the calendar ISO week of 2026-08-03-2026-08-09 (1 found). Showing the most recent 7 days with radar data instead: 2026-08-04-2026-08-10.
Covering Aug 04, 2026 to Aug 10, 2026 (UTC) · generated Aug 10, 2026.
This week’s radar flagged 0 hallucination-adjacent incidents across 0 named models (-3 vs. the previous 3-incident week). Severity split: 0 S1, 0 S2, 0 S3, 0 S4. By category: 0 hallucination, 0 jailbreak, 0 refusal, 0 bias. Verification funnel: 2 flagged by the heuristic → 0 verified by the LLM judge → 2 rejected as non-incidents (announcements, tutorials).
Severity breakdown
| Severity | Count | % of week |
|---|---|---|
| S1 | 0 | 0% |
| S2 | 0 | 0% |
| S3 | 0 | 0% |
| S4 | 0 | 0% |
Category breakdown
| Category | Count | % of week |
|---|---|---|
| Hallucination | 0 | 0% |
| Jailbreak | 0 | 0% |
| Refusal | 0 | 0% |
| Bias | 0 | 0% |
Top models by incidents
0 distinct model names were mentioned across this week’s incidents; the top 0 by incident count are ranked below.
No models with a named-model incident this week.
Most severe incidents this week
No incidents met the severity threshold this week.
Methodology
This is an MVP proxy index, not a verified incident registry. There is no dedicated hallucination-incident case database in GROUNDING yet — an “incident” here is any record from the AI-news radar’s own daily analysis journal (data/results/, ~2,500 records/week of AI news, papers, and community posts) that (1) names at least one model, (2) matches a hallucination / jailbreak / refusal / bias keyword pattern in its title, summary, or topics, and (3) was filed under a field-report category (model release, industry, or opinion coverage) rather than an academic research paper proposing a detection/mitigation method — the latter are excluded on purpose so this stays a field-incident signal, not a synthetic-benchmark leaderboard.
Consequences worth knowing before citing a number from this page: volume is low by construction (typically ~10-15 qualifying incidents/week); model names are raw NER extractions passed through a curated canonicalizer (known vendor-prefix/formatting duplicates are merged, e.g. “Claude Opus 4.8” and “Opus 4.8”), but genuinely ambiguous bare mentions spanning several concurrently-discussed versions (e.g. “Opus”) are left as their own entry rather than guessed onto one version — week-over-week deltas are still computed on the canonical string only; and “case links” point at the original source article, since no dedicated per-incident page exists yet. Severity S1-S4 is derived from the item’s global_importance score (1-5): 5→S1, 4→S2, 3→S3, 1-2→S4. Confidence (high/medium/low) reflects whether the keyword matched in the title, the prose, or only the tags.
Verification layer. Every candidate above is additionally run through an LLM judge (src.report.incident_verify, not a human reviewer) that decides whether it is a real field incident versus a research paper, announcement, or tutorial, and rechecks its category, severity, and named models. Candidates the judge rejects are removed from this report entirely; candidates it confirms are marked ”✓ Verified”, shown with the judge’s rechecked severity and category (which override the heuristic’s guess in every count above), and get a full write-up as a Case page under “/incidents/Cases/“. Everything else — not yet judged, or judged “uncertain” — is marked “Candidate”: still a keyword match, not yet independently verified by either the judge or a person.
See the current Hallucination Incident Index for the latest week.