Skip to content

Type: OpenAI model family entry

GPT-5 is an OpenAI model family entry tracked by GROUNDING for evaluations, product mentions, agent workflows, and builder-impacting changes.

Recent Updates

  • 2026-07-01: ArXiv paper evaluates how LLMs can induce belief states through planning and action (cs.CL updates on arXiv.org) · arxiv.orgAgents LLM Evals Gemini 2.5 Pro Claude 4
  • 2026-07-01: Paper on artificial swarm intelligence in large language models (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals OpenAI · Google · Anthropic · Gemini 2.5 Pro · Claude Sonnet 4.5
  • 2026-07-02: Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals
  • 2026-07-03: WAIC 2026 focuses on supernodes, optical interconnect, and compute infrastructure (量子位) · qbitai.comWAIC · Huawei Atlas 950 SuperPoD ZTE Xizhi Technology BiRen Technology MuXi Suiyuan Technology TianShu ZhiXin OEX dOCS FlagOS · Linux Eclipse · PyTorch Western Digital Hammerspace Shanghai Cube SuanFeng Information Lixun UniVista Yunhe Daoke Wuxin Qiong Fudan University Chuangzhi Academy Mohe Information Sugon scaleX scaleFabric H3C Shanghai David Patterson DeepSeek 671B
  • 2026-07-03: ProtoPilot and BioLab Bench aim to bring AI into real life science labs (量子位) · qbitai.comAgents Tool Use LLM Evals Code Agents Yongsheng Intelligent MGI Shanghai AI Laboratory OpenAI Ginkgo Bioworks · Google · Anthropic FutureHouse OpenTrons Huang Renxun · GPT-5.6 Sol · GPT-Rosalind
  • 2026-07-07: AgentGym2 benchmarks LLM agents in de-idealized real-world environments (cs.AI updates on arXiv.org) · arxiv.orgAgents Tool Use LLM Evals Gemini
  • 2026-07-07: RetroCoT: a forensic-reconstruction prompt that exposes framing-sensitive safety behavior (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals OpenAI · GPT-4o · gpt-4o-mini · GPT-5.4-mini
  • 2026-07-07: Replication Study Finds Natural Language Tools Improve Agent Tool-Calling Reliability (cs.CL updates on arXiv.org) · arxiv.orgTool Use Agents LLM Evals Johnson Martinez Chen Zhang Gemini 2.5 Pro
  • 2026-07-08: In-process retrieval may make per-step agent memory practical (cs.AI updates on arXiv.org) · arxiv.orgAgent Memory Agents arXiv · gpt-5-nano · GPT-5-mini
  • 2026-07-09: LiveOIBench introduces a competitive programming benchmark for evaluating LLMs (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals GPT-OSS 120B
  • 2026-07-09: Paper proposes deployment simulation to predict post-release LLM misbehavior (cs.LG updates on arXiv.org) · arxiv.orgLLM Evals Tool Use GPT-5.4
  • 2026-07-11: AegisDx proposes a safety-oriented framework for AI-assisted differential diagnosis (cs.AI updates on arXiv.org) · arxiv.orgAgents Tool Use RAG Context Engineering Yale New Haven Health System GPT-OSS 120B
  • 2026-07-12: AI Models Reach Human Diagnostic Accuracy But Lack True Clinical Reasoning, Say 2026 Studies (Все статьи подряд / Искусственный интеллект / Хабр) · habr.comLLM Evals OpenAI · Anthropic · Google Icahn School of Medicine Mikael Tordymann o1-preview · Claude-4.5-Opus Gemini 3.0 · Grok 4 Gemini 1.5 Flash
  • 2026-07-13: Geopolitical endorsement changes LLM policy judgments (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals OpenAI · Anthropic · Google · DeepSeek · Claude Sonnet · Gemini
  • 2026-07-13: OpenAI safety chief leaves as GPT-5.6 rolls out and safety team is reorganized (量子位) · qbitai.comAgents Code Agents Tool Use OpenAI · Wired · QbitAI · 量子位 Johannes Heidecke Ilya Sutskever Jan Leike 翁荔 Joshua Achiam Fidji Simo Saachi Jain Mia Glaese Mark Chen 听雨 · GPT-5.6 · GPT 5.5
  • 2026-07-14: UNIBROWSE proposes a data pipeline for multimodal BrowseComp agents (cs.CL updates on arXiv.org) · arxiv.orgAgents Tool Use LLM Evals Qwen3.5-35B-A3B · Gemini 2.5 Pro · Gemini 2.5 Flash
  • 2026-07-14: Claude Fable 5 evaluation on biomedical benchmarks highlights refusal behavior (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals Anthropic · Claude Fable 5
  • 2026-07-14: Tensor Is the Might (Hacker News) · zserge.com — Bellard
  • 2026-07-15: Frontier LLMs Lose Misconception Correction in Multi-Turn Medical Dialogues (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals AskDocs arXiv · Claude Haiku
  • 2026-07-16: Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing (cs.AI updates on arXiv.org) · arxiv.orgAgents Code Agents LLM Evals GPT-5.2 · GPT 5.5
  • 2026-07-16: Post-training alignment methods for biomedical data-to-text generation in small language models (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals openFDA Qwen
  • 2026-07-20: Import AI 465: open-weight cyber gaps, Kimi K3, and Demis’ policy plan (Import AI) · importai.substack.comLLM Evals Import AI UK government AI Security Institute AISI · DeepSeek · Kimi Demis · GLM-5.2 · DeepSeek V4 Pro · Claude Opus 4.6 Opus 4.5 · Claude Opus 4.5 · Sonnet 4.5 · Kimi K3
  • 2026-07-21: PlanFlip: Planning-Phase Prompt Injection Against Multi-Agent LLM Systems (cs.AI updates on arXiv.org) · arxiv.orgAgents Tool Use Context Engineering LLM Evals GPT-4o · Llama-3.3-70b · DeepSeek R1
  • 2026-07-22: MCP Workflows for QA Testing (Все статьи подряд / Искусственный интеллект / Хабр) · habr.comMCP Agents Code Agents Tool Use Anthropic · Claude · Sonnet 4.6
  • 2026-07-24: Rushes benchmark studies personalized engagement choices in interactive narratives (cs.CL updates on arXiv.org) · github.comLLM Evals Microsoft

FAQ

What is GPT-5?

GPT-5 is an OpenAI model family entry tracked by GROUNDING for evaluations, product mentions, agent workflows, and builder-impacting changes.

What does this page track?

Dated radar mentions, source links, related concepts, and builder-relevant context for GPT-5, collected automatically by GROUNDING.

When was GPT-5 last mentioned?

GPT-5 was most recently mentioned in a radar update dated 2026-07-24.