Type: AI model
GPT-5.4 is an OpenAI large language model released March 2026, initially in Thinking and Pro variants, with mini and nano versions following shortly after. The mini variant reached free-tier users while nano was offered via API. GROUNDING tracks GPT-5.4’s variants, availability tiers, and benchmark positioning.
Recent Updates
- 2026-06-29: Using visual app flows to constrain AI coding and cut token waste (Все статьи подряд / Искусственный интеллект / Хабр) · habr.com — Code Agents Context Engineering TRAE Figma · MiniMax · DeepSeek · GPT-5.2 Dola-seed 2.0 code · MiniMax-M2.7 · MiniMax M3 · kimi-k2.5 · DeepSeek-V3.2 Gemini 2.5flash 3Flash preview 3.1 Pro preview
- 2026-07-01: PolicyGuard uses dialogue context to verify policy adherence in LLM agents (cs.AI updates on arXiv.org) · arxiv.org — Agents Tool Use Context Engineering LLM Evals Claude Sonnet 4.6 · Gemini 2.5 Pro
- 2026-07-01: AutoTrainess packages workflows for autonomous LM post-training (cs.CL updates on arXiv.org) · arxiv.org — Agents Tool Use LLM Evals Context Engineering DeepSeek V4 Flash
- 2026-07-02: Senior SWE-Bench introduces a benchmark for evaluating agents like senior engineers (Hacker News) · senior-swe-bench.snorkel.ai — LLM Evals Code Agents mini-swe-agent Claude Opus 4.8 · Claude Sonnet 5 · GPT 5.5 · Claude Opus 4.7 · GLM-5.2 · kimi-k2.6 · Claude Sonnet 4.6 · Gemini 3.1 Pro · Gemini 3.5 Flash
- 2026-07-03: Quantifying the Affective Gap: Zero-Shot Emotion Classification Across Claude, GPT-5.4, and Gemini (cs.CL updates on arXiv.org) · arxiv.org — LLM Evals Anthropic · OpenAI · Google boltuix · claude-sonnet-4-6 · Gemini 2.5 Flash
- 2026-07-03: Benchmarking frontier LLMs on Arabic cultural and sociolinguistic knowledge (cs.CL updates on arXiv.org) · arxiv.org — LLM Evals Ghassan Al-Sumaidaee
- 2026-07-03: Rubric-based comparison of frontier models on clinician-authored reasoning tasks (cs.AI updates on arXiv.org) · arxiv.org — LLM Evals Claude Opus 4.7 · Gemini 3.1 Pro
- 2026-07-03: MultAttnAttrib proposes training-free multimodal attribution for long-document QA (cs.CL updates on arXiv.org) · arxiv.org — LLM Evals
- 2026-07-04: GPT-5.5 Codex token telemetry shows fixed reasoning-token spikes at 516, 1034, and 1552 (Hacker News) · github.com — LLM Evals OpenAI · GPT 5.5 · GPT-5.2
- 2026-07-05: A Comparative Experiment on Refactoring a LangGraph God Node with 11 Models (Все статьи подряд / Искусственный интеллект / Хабр) · habr.com — Agents Code Agents Data Sanity Habr · GPT 5.5 DeepSeek 4 Pro · Gemini 3.1 Pro · GLM-5.1 · Kimi-2.6 MiMo-2.5-pro · Opus 4.7 · Qwen 3.6 Plus · Qwen 3.7 Max · Fable 5
- 2026-07-07: Benchmarking rule adherence in semi-open textual sandboxes (cs.CL updates on arXiv.org) · arxiv.org — LLM Evals Claude Sonnet 4.6 · Gemini 3.5 Flash
- 2026-07-09: Paper proposes deployment simulation to predict post-release LLM misbehavior (cs.LG updates on arXiv.org) · arxiv.org — LLM Evals Tool Use GPT-5
- 2026-07-13: Vibe coding used to debug a Qlik Sense variable bug and build an audit script (Все статьи подряд / Искусственный интеллект / Хабр) · habr.com — Tool Use Qlik DAR KORUS Consulting QsAppMetadataConnector Chrome WebSocket Connectivity Tester Ilya Kerbatov
- 2026-07-14: PHITSBench benchmarks natural-language generation of PHITS inputs with execution-scored tasks (cs.AI updates on arXiv.org) · arxiv.org — LLM Evals Agents
- 2026-07-14: EvoClawBench tests whether agents can turn their own runs into reusable skills (cs.LG updates on arXiv.org) · arxiv.org — Agents LLM Evals Tool Use arXiv · MiniMax-M2.7 · DeepSeek V4 Pro
- 2026-07-14: Search APIs as decision surfaces for tool-using agents (cs.CL updates on arXiv.org) · arxiv.org — Agents Tool Use RAG LLM Evals Brave · Tavily · Firecrawl · kimi-k2.6
- 2026-07-15: Anthropic Research: Four New Agentic Misalignment Failure Modes in Frontier Models (Anthropic) · alignment.anthropic.com — Agents Anthropic · OpenAI · Google DeepMind · xAI · DeepSeek · Moonshot AI MJ Rathbun · Claude Opus 4.8 · Claude Opus 4.7 · Claude Opus 4.6 · Claude Opus 4.5 · Claude Sonnet 4.6 · Claude Mythos Preview · GPT 5.5 · Gemini-3.1-Pro · Gemini 3 Flash · Gemini 3.5 Flash · Grok 4.3 · DeepSeek V4 · kimi-k2.6
- 2026-07-16: PM-Bench: Evaluating Prospective Memory in LLM Agents (cs.AI updates on arXiv.org) · arxiv.org — Agent Memory Agents LLM Evals
- 2026-07-16: Rethinking the Evaluation of Harness Evolution for Agents (cs.AI updates on arXiv.org) · arxiv.org — Agents LLM Evals OpenAI · Anthropic · Claude Opus 4.6
- 2026-07-16: When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation (cs.CL updates on arXiv.org) · arxiv.org — RAG LLM Evals Faizan Faisal DeepSeek-V4-Flash · Gemma 4 E4B
- 2026-07-16: STOCKTAKE: A benchmark for separating perception from action in LLM agents (cs.AI updates on arXiv.org) · arxiv.org — Agents LLM Evals Claude Sonnet 5 · DeepSeek V4 Pro · Grok 4.5
- 2026-07-20: ARC-AGI-3 study isolates execution, simplification, and verification in Codex-based agents (cs.AI updates on arXiv.org) · arxiv.org — Code Agents LLM Evals Agents GPT 5.5 · GPT-5.6 Sol
- 2026-07-23: OpenAI’s failed security test became a case study in agentic exploit capability (Simon Willison’s Weblog) · simonwillison.net — Agents LLM Evals OpenAI · Hugging Face · Anthropic · Google · UC Berkeley Max Planck Institute UC Santa Barbara Arizona State Simon Willison · Claude Mythos Preview · GPT 5.5 · Claude Opus 4.7 · Claude Opus 4.6 · Gemini-3.1-Pro
- 2026-07-23: OpenAI’s security test allegedly escaped its sandbox and attacked Hugging Face (Hacker News) · simonwillison.net — Agents LLM Evals Tool Use OpenAI · Hugging Face · Anthropic · Google · UC Berkeley Max Planck Institute UC Santa Barbara Arizona State · Claude Mythos Preview · GPT 5.5 · Claude Opus 4.7 · Claude Opus 4.6 · Gemini-3.1-Pro
- 2026-07-27: Claude models benchmarked on incremental coding: Opus 5 reaches 24% on SlopCodeBench (Hacker News) · github.com — LLM Evals Code Agents OpenAI University of Wisconsin Madison GOrlanski · Opus 5 · Opus 4.8 · Opus 4.6 · Sonnet 5 · Fable 5.6 Sol
FAQ
What is GPT-5.4?
GPT-5.4 is an OpenAI large language model released March 2026, initially in Thinking and Pro variants, with mini and nano versions following shortly after. The mini variant reached free-tier users while nano was offered via API. GROUNDING tracks GPT-5.4’s variants, availability tiers, and benchmark positioning.
What does this page track?
Dated radar mentions, source links, related concepts, and builder-relevant context for GPT-5.4, collected automatically by GROUNDING.
When was GPT-5.4 last mentioned?
GPT-5.4 was most recently mentioned in a radar update dated 2026-07-27.
Category: Text / Language Models