Skip to content

Type: AI model or model family

GRPO appears in the radar stream as a model or model family. This page is a living index of dated mentions and sources — open Recent Updates and Backlinks for context, not a full product brief.

Recent Updates

  • 2026-07-01: ECHO proposes selective turn memory for long-horizon agent RL (cs.LG updates on arXiv.org) · arxiv.orgAgents Agent Memory Context Engineering LLM Evals arXiv SUPO ECHO
  • 2026-07-03: OpenReward trains a tool-augmented reward model for long-form agentic tasks (cs.CL updates on arXiv.org) · arxiv.orgTool Use LLM Evals Ziyou Hu OpenRM LLMs
  • 2026-07-03: Input rewriting is not reliably enough for dialogue discourse parsing (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals SDRT
  • 2026-07-09: Single-Rollout Asynchronous Optimization for agentic RL (cs.LG updates on arXiv.org) · arxiv.orgAgents Code Agents LLM Evals GLM · GLM-5.2
  • 2026-07-20: Cotype describes how it trains compact tool-using agents from real execution traces (Все статьи подряд / Искусственный интеллект / Хабр) · habr.comAgents Tool Use LLM Evals Cotype τ²-bench verl-agent gym Cotype Pro 3 Cotype Light 3
  • 2026-07-24: CORE: Contrastive Reflection for faster reasoning improvement (alphaXiv) — LLM Evals RAG Linas Nasvytis Simon Jerome Han Ben Prystawski Satchel Grant Noah D. Goodman Judith E. Fan GEPA MemRL

FAQ

What is GRPO?

GRPO appears in the radar stream as a model or model family. This page is a living index of dated mentions and sources — open Recent Updates and Backlinks for context, not a full product brief.

What does this page track?

Dated radar mentions, source links, related concepts, and builder-relevant context for GRPO, collected automatically by GROUNDING.

When was GRPO last mentioned?

GRPO was most recently mentioned in a radar update dated 2026-07-24.