🛰 AI Brief — 23 June 2026
🥇 PlanBench-XL benchmarks long-horizon tool-use planning in large tool ecosystems ·
prio 13This is directly relevant to builders working on agentic workflows because it evaluates a concrete failure mode: agents operating across large tool sets with limited visibility and runtime disruption. The benchmark and reported drop in performance give a useful signal that tool discovery and recovery behavior are still brittle in real agent systems. arxiv.org · 48 sources · Agents Tool Use LLM Evals GPT-5.4
🥈 EvoEmbedding for long-context retrieval and agentic memory ·
prio 13This is directly relevant to builders working on retrieval and agent workflows because it proposes an embedding approach that changes representation as context evolves, rather than treating text chunks as static. The paper also claims the model can improve a naive RAG pipeline enough to beat dedicated agentic memory systems, which is a meaningful signal for people evaluating memory and retrieval architectures. arxiv.org · Embeddings Long Context Agent Memory RAG Agents EvoEmbedding Qwen3-Embedding-8B KaLM-Embedding-Gemma3-12B
🥉 CORE compresses RAG prompts for edge QA without using auxiliary small models ·
prio 13This is directly relevant to builders working with RAG and prompt assembly because it targets redundant retrieved context, which is a common source of wasted tokens and latency. The paper is especially useful for teams that care about running QA systems on constrained hardware, since it focuses on memory, speed, and energy tradeoffs on edge devices. arxiv.org · 2 sources · RAG Context Engineering NVIDIA Huawei arXiv
4️⃣ AgentRiskBOM proposes a machine-readable security BOM for agentic AI systems ·
prio 13The paper is directly relevant to builders working on agents because it proposes a structured way to describe what an agent can access, change, delegate, and do externally before incidents happen. For AI-builder teams, the practical value is in the schema and evaluation artifacts, which focus attention on authority, auditability, and deployment drift rather than vague capability claims. arxiv.org · Agents Tool Use RAG Code Agents
5️⃣ Managing a 114k-line browser game when the code no longer fits in context ·
prio 12The post gives a concrete example of how an AI-assisted coding workflow can remain workable on a very large codebase instead of collapsing once the context window is exceeded. For builders working with coding agents, it highlights practical scaffolding such as planning files, a source-of-truth document, dependency-aware edits, and a custom MCP helper for accessing large files. habr.com · Context Engineering MCP Code Agents Habr
⚠️ Knowledge Gaps
🚀 Models & Releases (1)
7Unlimited-OCR is released with paper, model, and SGLang inference instructions · github.com · ModelScope Hugging Face NVIDIA Baidu PyMuPDF
🧪 Research Papers (121)
12Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning · arxiv.org · Agent Memory Context Engineering12RAG compression changes reader scaling and model rankings · arxiv.org · RAG RAG Evaluation LLM Evals Qwen-7B GPT-4.1-mini11Hierarchical subgoals improve how agents read demonstrations · arxiv.org · Agents Context Engineering11RIZZ proposes verifier-gated continual adaptation for black-box agents · arxiv.org · Agents Agent Memory Context Engineering11Don’t Blindly Trust It: How Unreliable Feedback Breaks Tool-Using LLM Agents · arxiv.org · Agents Tool Use LLM Evals Qwen2.5-7B11Negative Knowledge as Shared Memory for AutoResearch · arxiv.org · Agent Memory Agents11AdaMem proposes feedback-driven memory policies for long-horizon LLM agents · arxiv.org · Agent Memory LLM Evals arXiv Mem0 alphaXiv11CELEUS proposes anytime-valid confidence intervals for LLM evaluation · arxiv.org · LLM Evals10CFAgentBench benchmarks autonomous construction-finance agents in a self-hostable environment · arxiv.org · Agents LLM Evals10Power Systems Agent Benchmark introduces executable evaluation for power-engineering agents · arxiv.org · Agents LLM Evals10DISC: Structured verification loops for LLM self-correction · arxiv.org · LLM Evals Sonnet 4.510Skill Coverage Measures How Thoroughly Agent Skills Are Tested · arxiv.org · LLM Evals10Scientific fine-tuning is linked to worse factuality in a multi-domain LLM evaluation · arxiv.org · LLM Evals10Coherence Under Commitment: An Evaluation for Vacuous Abstention in LLM Reasoning · arxiv.org · LLM Evals Qwen2.5-3B TinyLlama-1.1B10Nous: A Predictive World Model for Long-Term Agent Memory · arxiv.org · Agent Memory LLM Evals gpt-4o-mini10GRAIDES proposes a lightweight schema for centralizing generative AI evaluation data · arxiv.org · LLM Evals Westminster City Council10ChainWorld builds long-horizon desktop agent evaluations from atomic OSWorld tasks · arxiv.org · Agents LLM Evals10FairTutor: Equity-Aware Pedagogical LLM Routing for Budget-Constrained AI Tutoring · arxiv.org · Agents Tool Use LLM Evals10Calibration Is Not Control: Oversight for LLM Agents Should Use Intervention Value · arxiv.org · Agents LLM Evals10LLM self-training can improve fast, then collapse within the same run · arxiv.org · LLM Evals Qwen-2.5 3B Qwen 2.5 7B Gemma-3-4B10PeerCheck studies how CoT and RAG affect LLM-generated academic reviews · arxiv.org · RAG LLM Evals10Benchmarking Small Language Models for Arabic NLP · arxiv.org · LLM Evals Gemma 3 Aya C4AI Command Arabic GPT-4.1-mini10Storyline Trees for Long-Form Narrative QA · arxiv.org · RAG Long Context Chunking arXiv10PrivacyAlign: Human-annotated privacy alignment for LLM agents · arxiv.org · Agents LLM Evals10Paper argues that LLM user simulators miss real buyer disengagement · arxiv.org · LLM Evals DeepSeek10Test-Time Steering for Temporal Fact Conflicts in Open-Weight LLMs · arxiv.org · LLM Evals Open Source LLMs Qwen-2.5-1.5B Qwen 2.5 7B Mistral-7B-v0.39Prompt Injection as Role Confusion · simonwillison.net · gpt-oss-20b9Reference-free physical consistency evaluation for generated videos · arxiv.org · LLM Evals9Pre-generation hallucination detection via soft-target attention probing · arxiv.org · LLM Evals9Study compares LLM translation quality and metric reliability for Hausa and Fongbe · arxiv.org · LLM Evals arXiv gpt-4o-mini Claude Sonnet 4 Gemini 2.5 Flash9TALAS proposes a more efficient way to distill sentence embeddings · arxiv.org · Embeddings9CulMind introduces a benchmark for multimodal reasoning in Chinese cultural heritage · arxiv.org · LLM Evals arXiv9Design-Time Verification for Agentic AI Workflows · arxiv.org · Agents Tool Use9DEMM-Bench benchmarks governance evidence sufficiency for agent runtimes · arxiv.org · LLM Evals Agents9NL2Scratch introduces an executable benchmark and semantic metric for Scratch generation · arxiv.org · LLM Evals9Benchmarking self-awareness in LLM agents · arxiv.org · Agents Tool Use LLM Evals9Paper argues repeated DPO can cause ‘scientific amnesia’ in continual post-training pipelines · arxiv.org · LLM Evals Qwen2.5-7B-Instruct9Discretizing Reward Models · arxiv.org · LLM Evals9A budget-matched study finds LLM advice does not improve tabular hyperparameter optimization · arxiv.org · LLM Evals9The Reversal Curse in LLMs · arxiv.org · LLM Evals arXiv GPT-3 Llama-1 ChatGPT8CURIOBOT studies how curiosity-oriented tutoring language changes exploratory behavior · arxiv.org · LLM Evals arXiv8Hypothesis-Driven Skill Optimization for LLM Agents · arxiv.org · Agents Tool Use Context Engineering Qwen3-8B Qwen3.6-27B8ARTS uses a reasoning LM to guide tree search for automated scientific discovery · arxiv.org · Agents Context Engineering LLM Evals Qwen3-4B Gemini 3 Pro8TSCognition benchmark and TSAlign framework for LLM time series reasoning · arxiv.org · LLM Evals8Experimental study of how frontier AI systems scale on Project Euler math problems · arxiv.org · LLM Evals8When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study · arxiv.org · arXiv8LLMs show ordering preferences in binomial pairs, but not full corpus distribution matching · arxiv.org · LLM Evals8Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Quality Evaluation · arxiv.org · LLM Evals8IPA-based multilingual tokenization for better cross-language consistency · arxiv.org · arXiv8MedHal-Loc benchmarks localization faithfulness in medical hallucination detectors · arxiv.org · LLM Evals8PuMVR benchmarks script bias in multilingual VLMs · arxiv.org · LLM Evals8A-Evolve-Training reports autonomous post-training of a 30B Nemotron · arxiv.org · Agents LLM Evals NVIDIA Nemotron 30B Nemotron8Test-Time Training with Next-Token Prediction · arxiv.org · Long Context arXiv Llama-3.1-8B Mistral-7B-v0.3 Qwen38ForEx verifies LLM explanations for logical fallacy detection in Lean4 · arxiv.org · LLM Evals8HALO trains an LLM orchestrator for verified PDDL planning · arxiv.org · Agents LLM Evals GPT-5-mini Gemini 3 Flash8Paper on pedagogically aligned LLM tutors for math mistake remediation · arxiv.org · LLM Evals8CalVerT adds calibrated confidence and grounding signals to QA agents · arxiv.org · Agents RAG8PulseCX: Breaking the Closed-World Assumption in Real-Time CX · arxiv.org · Agents Agent Memory8Paper Proposes Hardness Adjusted Transfer Score for Cross-Lingual Evaluation · arxiv.org · LLM Evals arXiv8Metanym Game proposes a self-contained peer benchmark for LLM structural intelligence · arxiv.org · LLM Evals8Latent Personal Memory turns user history into dynamic soft prompts · arxiv.org · Agent Memory Context Engineering arXiv Qwen3-1.7B Qwen3-4B8LaBSE and curriculum learning for multilingual polarization detection · arxiv.org · Embeddings RAG Context Engineering LaBSE Qwen8Decodable but Not Faithful: Verifier-Coupled Reasoning for Rationale Traces · arxiv.org · LLM Evals arXiv8HERALD: CPU-GPU Cooperative KV Cache Retrieval for Block Diffusion LLM Serving · arxiv.org · Long Context Context Engineering8Sakana Fugu technical report on orchestrator models for agent teams · arxiv.org · Agents Tool Use Sakana Fugu Fugu-Ultra8Paper argues coding agents can replace mandatory human code review · arxiv.org · Code Agents Agents8ParallelKernelBench benchmarks where LLMs break on multi-GPU kernel generation · twitter.com · Together Compute alphaXiv7Measured coupling gain and validity diagnostics for LLM agent societies · arxiv.org · Agents LLM Evals7Many-shot in-context learning for NER can match supervised baselines and support low-resource annotation · arxiv.org · Context Engineering BERT7Clinical term extraction from ALS notes with open-source small models · arxiv.org · LLM Evals arXiv Qwen3-4B-Instruct-2507 Hammer2.1-7b7Per-Entity Bias Mapping study on AI visibility and fabricated citations · arxiv.org · RAG arXiv7Score granularity in black-box LLM classification · arxiv.org · LLM Evals7Holmes: Multimodal Agentic Diagnosis for Mixed-Language Mobile Crashes at Industrial Scale · arxiv.org · Agents Code Agents RAG WeChat7FirstPass: a peer-review dataset and fine-tuned model built from multi-round editorial outcomes · arxiv.org · LLM Evals Nature Communications arXiv Qwen2.5-7B-Instruct Gemini-3.1-flash-lite-preview7Investigating Linguistic Steering in Large Language Models · arxiv.org · o3 gpt-4o-mini Phi-3 Llama 3 70B DeepSeek R17Condition-Aware Analysis of Chain-of-Thought Distillation · arxiv.org · LLM Evals7ARCO proposes co-evolving rubrics and step-level rewards for multi-step LLM agents · arxiv.org · Agents7Darwin Mobile Agent proposes a roadmap for self-evolving GUI agents · arxiv.org · Agents Agent Memory7Benchmarking LLMs for Japanese grapheme-to-phoneme conversion · arxiv.org · LLM Evals arXiv alphaXiv CatalyzeX DagsHub7AutoRAS: A framework for designing robust agentic systems with symbolic primitives · arxiv.org · Agents Tool Use LLM Evals7Post-training recipe matters more than model family for multi-agent conversational behavior · arxiv.org · Agents LLaMA Qwen7Inverse Turing Bench evaluates models on human vs. AI dialogue detection · arxiv.org · LLM Evals GPTZero Anthropic OpenAI Claude Opus 4.67Tree-of-Thought methods show opposite failure modes across compute budgets · arxiv.org · LLM Evals Llama-3B Llama-8B7Synthetic Audio Pipeline for Air Traffic Control ASR · arxiv.org · Whisper7Measuring multilingual LLM inference energy across languages · arxiv.org · LLM Evals7HEAL targets reproducible LLM inference under 16-bit precision · arxiv.org · LLM Evals7Mechanistic study finds retrieval heads matter for long-context recall, and RoPE does not stop them from forming · arxiv.org · Long Context arXiv OLMo-2 Qwen Llama 3.16ORBIT proposes training-free multi-attribute steering for language models · arxiv.org · LLM Evals Llama-3.2-3B Qwen 2.5 7B Llama-3.1-8B6ARIA proposes causal-aware LLM reasoning for materials discovery · arxiv.org · RAG LLM Evals6Geometry-Aware Online Scheduling for LLM Serving Proposes SVF and 1-bit SVF · arxiv.org · arXiv vLLM Llama 3.16Reasoning models show only modest ability to detect CoT edits · arxiv.org · LLM Evals6Context Drift in Multi-Agent LLM Systems · arxiv.org · Agents Tool Use LLM Evals Claude Haiku6Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents · arxiv.org · DINOv26A Framework for Evaluating CEFR-Controlled Arabic Text Generation · arxiv.org · LLM Evals Taha-196Process-Reward Tactic Evolution for long-horizon Galaxy workflow execution · arxiv.org · Agents Tool Use Galaxy6Document-tuned transformer representations improve person-level mental health assessment · arxiv.org · Embeddings6DrugBench evaluates AI control protocols for medication safety in medical QA · arxiv.org · LLM Evals FDA arXiv6SCOPE proposes conformal probing for OOD rejection in LLM services · arxiv.org · LLM Evals6Hypothesis-Disciplined Multi-Agent Automated Formalization of Asymptotic Statistical Theory · arxiv.org · Agents6Rational value risk in LLM reasoning · arxiv.org · LLM Evals Llama 3.1 Qwen 2.5 Tulu-3 GPT-5.26Answer Engineering uses local trajectory editing to improve protocol adherence in LLM outputs · arxiv.org · LLM Evals arXiv6SPARC: A Multi-Agent System for Electrical Circuit Question Answering · arxiv.org · Agents Tool Use6CAT-Translate: Compact Open-Source Models for Japanese-English Translation · arxiv.org · LLM Evals Open Source LLMs6Confident layer decoding reduces final-layer alignment perturbations · arxiv.org · LLM Evals6Keyless Attention removes key projections and cuts KV cache usage by 50% · arxiv.org · LLM Evals GPT-2 280M GPT-2 557M Pythia 410M Qwen2 1.5B6Token-level comparison of transformers and hybrid language models · arxiv.org · LLM Evals OLMo 3 Olmo Hybrid6Measuring value-structure alignment in LLMs with symmetric Q-sorts · arxiv.org · LLM Evals arXiv6Factual retrieval in LLMs appears distributed, redundant, and non-contiguous · arxiv.org · arXiv arXivLabs alphaXiv CatalyzeX DagsHub6Scalable hierarchical attention for multi-turn jailbreak detection · arxiv.org · LLM Evals Long Context Claude Opus 4.76Using LLM internal artifacts to flag wrong legal classification outputs · arxiv.org · LLM Evals6PEAR proposes adaptive routing for multi-agent debate · arxiv.org · Agents6A Causal DAG Prior for Synthetic Time-Series Classification Datasets · arxiv.org · LLM Evals arXiv alphaXiv CatalyzeX DagsHub6Fast-TurboQuant proposes multiplier-free online vector quantization · arxiv.org · RAG Embeddings OpenAI DBpedia OpenAI-3 Large6Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs · arxiv.org · LLM Evals MR-LoRA RePRo MLLM-CL6Study Finds No Single Loss-Optimizer Pair Works Best Across a Heterogeneous Model Pool · arxiv.org6ELADO benchmarks failure modes in neural operator learning for elliptic PDEs · arxiv.org · LLM Evals6Empirical study maps OpenPangu quantization behavior on Huawei Ascend NPUs · arxiv.org · LLM Evals Huawei OpenPangu OpenPangu 1B OpenPangu 7B6Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective · arxiv.org · arXiv6Understanding Latent Flow Models for Tabular Data Synthesis: Targets, Paths, and Sampling · arxiv.org6Pact: Anonymous Credentials for the Web · hacks.mozilla.org · Agents Distilled Google5AlphaMemo proposes structured search-process memory for alpha mining agents · arxiv.org · Agents Agent Memory LLM Evals
🛠 Tools & Frameworks (7)
8CuratorKIT: Open-Source Pipeline for LLM Post-Training Data Curation · arxiv.org · arXiv LiteLLM Hugging Face8KiSinWi: an AutoML platform built around multi-agent workflows · habr.com · Agents Tool Use LLM Evals8Bun-sqlgen adds type-safe raw SQL for Bun · github.com · Bun7alphaXiv launches a compute catalog for low-cost on-demand GPUs · openresearch.sh · alphaXiv openresearch.sh7Datasette 1.0a35 adds create and alter table actions plus template context docs · simonwillison.net · Context Engineering6Mistral OCR 4 adds structured document output, self-hosting, and multilingual OCR · mistral.ai · RAG Mistral Mistral OCR 46Modal launches Auto Endpoints for self-serve LLM inference · modal.com · Tool Use Modal Cognition Decagon Fathom
💬 Opinions (14)
12How a simple local LLM call becomes a small FastAPI RAG backend · habr.com · RAG Chunking Embeddings Vector Database Ollama11Building an AI Assistant Over a Company Codebase: A Zvuk Team Case Study · habr.com · RAG Codebase Indexing Context Engineering GIGASCHOOL Zvuk10Why GenAI assistants need platform logic beyond retrieval · habr.com · RAG LLM Evals Context Engineering9Corporate AI only works when the knowledge base is current · habr.com · RAG TEAMLY QSOFT Yandex Severstal7OpenAI Candidate Shares Her AI Interview Prep Notes · qbitai.com · OpenAI Google NVIDIA Stanford The University of Washington7Poll on How Critical Full Structured Output Is for Local Models · developers.openai.com · OpenAI BitGN llguidance7A Claude Code case study on anomaly hunting in school olympiad results · habr.com · Code Agents Chess.com Habr7Karpathy on Claude as a persistent team member inside org workflows · twitter.com · Agents Tool Use Context Engineering6Benchmarking Mythos Against Blind Bug-Finding Runs · swelljoe.com · LLM Evals Code Agents Anthropic Mythos Opus6Prototype-first AI dev workflow with Claude and DeepSeek · t.me · Code Agents DeepSeek Claude6Why LLM-generated specs do not eliminate product ambiguity · t.me6Who Does What? Team Topologies for the Agentic Platform · blog.owulveryck.info · Agents6Why harness loops can make agent-written code harder to trust · lucumr.pocoo.org · Agents Code Agents Tool Use Context Engineering6Probing for an internal “thinking” direction in an LLM · habr.com
📦 Other (2)
6Book on Using LLMs for Data Analysis · habr.com · Agents OpenAI LangChain LlamaIndex Anthropic6OpenAI opens applications for DevDay 2026 in San Francisco · devday.openai.com · OpenAI
FAQ
What is in the 2026-06-23 AI brief?
The 2026-06-23 brief selected 150 signal items for AI builders and filtered 253 items as noise, using the radar’s community-relevance scoring.