Type: AI model or model family
GRPO appears in the radar stream as a model or model family. This page is a living index of dated mentions and sources — open Recent Updates and Backlinks for context, not a full product brief.
Recent Updates
- 2026-07-01: ECHO proposes selective turn memory for long-horizon agent RL (cs.LG updates on arXiv.org) · arxiv.org — Agents Agent Memory Context Engineering LLM Evals arXiv SUPO ECHO
- 2026-07-03: OpenReward trains a tool-augmented reward model for long-form agentic tasks (cs.CL updates on arXiv.org) · arxiv.org — Tool Use LLM Evals Ziyou Hu OpenRM LLMs
- 2026-07-03: Input rewriting is not reliably enough for dialogue discourse parsing (cs.CL updates on arXiv.org) · arxiv.org — LLM Evals SDRT
- 2026-07-09: Single-Rollout Asynchronous Optimization for agentic RL (cs.LG updates on arXiv.org) · arxiv.org — Agents Code Agents LLM Evals GLM · GLM-5.2
- 2026-07-20: Cotype describes how it trains compact tool-using agents from real execution traces (Все статьи подряд / Искусственный интеллект / Хабр) · habr.com — Agents Tool Use LLM Evals Cotype τ²-bench verl-agent gym Cotype Pro 3 Cotype Light 3
- 2026-07-24: CORE: Contrastive Reflection for faster reasoning improvement (alphaXiv) — LLM Evals RAG Linas Nasvytis Simon Jerome Han Ben Prystawski Satchel Grant Noah D. Goodman Judith E. Fan GEPA MemRL
FAQ
What is GRPO?
GRPO appears in the radar stream as a model or model family. This page is a living index of dated mentions and sources — open Recent Updates and Backlinks for context, not a full product brief.
What does this page track?
Dated radar mentions, source links, related concepts, and builder-relevant context for GRPO, collected automatically by GROUNDING.
When was GRPO last mentioned?
GRPO was most recently mentioned in a radar update dated 2026-07-24.
Category: Text / Language Models