Skip to content

This is GROUNDING’s LLMs topic hub: a running, dated log of developments the radar has mapped to LLMs, refreshed as new signal arrives.

Key Developments

  • 2026-07-24: LLMs Struggle When User Intent Changes Mid-Conversation (cs.LG updates on arXiv.org) · arxiv.org LLM Evals
  • 2026-07-24: Anthropic introduces Claude Opus 5 (ElKornacio) · anthropic.com LLM Evals Anthropic · Zapier · Claude Opus 5 · Claude Fable 5 · Mythos 5 · Opus 4.8 Claude Max Claude Pro
  • 2026-07-24: Claude Opus 5 launches with stronger coding and knowledge-work benchmark results (Hacker News) · anthropic.com LLM Evals Claude Opus 5 · Claude Fable 5 · Mythos 5 · Opus 4.8
  • 2026-07-24: CORE: Contrastive Reflection for faster reasoning improvement (alphaXiv) LLM Evals Linas Nasvytis Simon Jerome Han Ben Prystawski Satchel Grant Noah D. Goodman Judith E. Fan GRPO GEPA MemRL
  • 2026-07-24: Drone-Bench measures frontier models on simple drone surveillance tasks (Hacker News) · andonlabs.com LLM Evals
  • 2026-07-24: Opus 5 tops the Artificial Analysis Intelligence Leaderboard (Hacker News) · artificialanalysis.ai LLM Evals Long Context Artificial Analysis · Claude Opus 5 · Claude Fable 5 · GPT-5.6 Sol Mercury 2 HyperNova 60B 2605 Granite 4.0 H Small Gemma 3n E4B Instruct Nova Micro Sarvam 30B · Gemini 2.5 Flash Lite · Command A+ · Gemini 2.5 Flash · GLM-5.2 · MiniMax M3 · DeepSeek V4 Pro
  • 2026-07-25: A catalog of 3,607 reported AI agent misbehavior incidents (Hacker News) · rewardhacking.org LLM Evals GitHub · Hacker News LessWrong · X
  • 2026-07-25: Vivix launches A1, a real-time interactive multimodal model with a unified streaming architecture (量子位) · qbitai.com Context Engineering Vivix 量子位 · QbitAI henry A1 W1
  • 2026-07-25: UK AISI and CAISI publish a preliminary cyber-capability assessment of Kimi K3 (Hacker News) · nist.gov LLM Evals UK Artificial Intelligence Security Institute U.S. Center for AI Standards and Innovation Moonshot AI · Carnegie Mellon University · Kimi K3 · GLM-5.2
  • 2026-07-25: ARC-AGI-3 leaderboard compares performance against task cost (Hacker News) · arcprize.org LLM Evals Kaggle ARC Prize ARC-AGI-1 ARC-AGI-2 ARC-AGI-3 GPT-4.5 Claude 3.7 o1 pro · Gemini 3 Pro
  • 2026-07-25: A debate on whether agents are already useful or still just getting started (量子位) · qbitai.com Context Engineering Long Context QbitAI · Alibaba Kujing Technology Miaopai CCF · Tsinghua University Nanjing University Shanghai Jiao Tong University Tianjin University Shandong University Shanghai AI Laboratory Jin Lei Gao Yang Hao Jianyie Han Zhongyi Wen Ying Zhou Hao Zhang Hangfan Du Yanlong
  • 2026-07-25: Anthropic’s Opus 5 is reported to match or beat Fable 5 on several benchmarks while Claude Code’s system prompt was cut by more than 80% (量子位) · qbitai.com Context Engineering LLM Evals Anthropic · QbitAI · Cursor · OpenAI AlphaSchool 克雷西 am.will OmedTheVibeCoder Alex Ermolov Chetaslua Noema Matt Shumer Victor M Thariq · Opus 5 · Opus 4.8 · Fable 5 · GPT-5.6 Sol · GPT-5.6 · Kimi K3
  • 2026-07-25: Claude Code trims system prompts as context engineering shifts (Hacker News) · claude.com Context Engineering LLM Evals Claude Opus 5 · Claude Fable 5
  • 2026-07-26: A personal report on agentic coding, testing, and model variance (Hacker News) · danluu.com LLM Evals OpenAI · Mastodon · Playwright · GPT-5.0 · GPT-5.1
  • 2026-07-26: World-model-optimizer: Continuously optimize agents and reduce inference costs through model routing and distillation (Hacker News) · github.com LLM Evals experientiallabs OpenRouter · e2b
  • 2026-07-27: Wattage: A token-spend profiler and cost-regression gate for AI agents (Hacker News) · github.com Context Engineering faizannraza
  • 2026-07-27: Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA (cs.LG updates on arXiv.org) · arxiv.org Context Engineering
  • 2026-07-27: Lost in Context: Addressing Context Anxiety in Large Language Models (cs.AI updates on arXiv.org) · arxiv.org Context Engineering
  • 2026-07-27: The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents (cs.AI updates on arXiv.org) · arxiv.org Context Engineering LLM Evals
  • 2026-07-27: Kimi-K3: Open-Source Frontier Model with Native Agentic and Coding Capabilities (Hacker News) · huggingface.co Context Engineering Long Context Moonshot AI · Hugging Face · Kimi K3
  • 2026-07-27: Terence Tao warns mathematics faces century-long crisis as AI proves capable of solving research-level problems (量子位) · qbitai.com LLM Evals Terence Tao Wang Hong Deng Yu Jean Bourgain William Thurston
  • 2026-07-27: MirrorCode Benchmark: AI Systems Complete Complex Programming Tasks in Hours Instead of Weeks (Import AI) · importai.substack.com LLM Evals Epoch METR · Anthropic · Apple · Claude Opus 4.7 · GPT 5.5
  • 2026-07-28: How Context Attribution Handles What the Model Already Knows (cs.CL updates on arXiv.org) · arxiv.org Context Engineering
  • 2026-07-28: Rebuilding legacy sites using AI without reading the code (Все статьи подряд / Искусственный интеллект / Хабр) · habr.com Context Engineering
  • 2026-07-28: Small language models need fine-tuning to effectively use provided legal context (cs.CL updates on arXiv.org) · huggingface.co Context Engineering LLM Evals Qwen3.5

FAQ

What is the LLMs topic?

The LLMs topic is GROUNDING’s hub for LLMs: a running, dated log of developments, releases, and mentions the radar has tracked, refreshed as new signal arrives.

What does the LLMs topic page track?

Key developments the GROUNDING radar mapped to LLMs, updated through 2026-07-28.

How current is this page?

The most recent LLMs development listed here is dated 2026-07-28; the radar refreshes hourly.