🛰 AI Brief — Jul 14, 2026
How to read
prioand sources
prio Nis the radar’s practical-relevance score for this item (higher runs first; items at or below the noise threshold are filtered out as noise). Under each signal: Concepts / Entities are graph links; Source / N sources list every outbound link for that story.
🥇 CAFE turns compound AI evaluation into a factorial experiment ·
prio 13This gives builders a concrete way to measure which part of a compound AI pipeline is actually driving quality, instead of only searching for a better overall configuration. It is especially relevant for retrieval-augmented QA and for teams that need to compare quality against cost and latency with statistical support. Concepts: LLM Evals RAG Evaluation RAG Source: arxiv.org
🥈 Prompt wrapper formatting can materially change benchmark results, according to a new arXiv paper ·
prio 12Builders who compare models or rely on structured outputs, because the paper says wrapper formatting and parseability can materially move scores. It adds a concrete warning that benchmark numbers can be fragile unless wrapper variance and compliance are reported alongside accuracy. Concepts: LLM Evals Tool Use Entities: OpenRouter 17 sources: arxiv.org, habr.com, arxiv.org, contextvault.dev, [sleuth-io.github.io](https://sleuth-io.github.io/sx/2026/07/10/the community’s-dropbox-is-now-a-skill-server.html), habr.com, habr.com, minor.gripe, arxiv.org, habr.com, qbitai.com, arxiv.org, developers.googleblog.com, latent.space, qbitai.com, habr.com, qbitai.com
🥉 ANCHOR audits CLI agents against persistent malicious users ·
prio 12For builders of agents and CLI workflows, this is a concrete warning that single-turn refusals are not enough to characterize safety. The paper frames evaluation around persistent, adaptive misuse rather than isolated prompts, which is directly relevant to anyone building or auditing tool-using agents. Concepts: Agents Tool Use LLM Evals Code Agents Entities: arXiv 12 sources: arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org, arxiv.org
4️⃣ GRASP trains agentic RAG policies to choose between semantic search, keyword search, and paragraph reading ·
prio 12The paper is directly relevant to builders working on agentic retrieval systems because it studies when to retrieve, which retrieval mode to use, and how much context to expand. Its main practical signal is that learned coordination across search modes and context granularity can improve both retrieval quality and answer performance on multi-hop tasks. Concepts: RAG Agents Tool Use Hybrid Search Context Engineering Source: arxiv.org
5️⃣ DoorDash describes an LLM-driven metadata pipeline for food catalogs ·
prio 12This is a concrete production example of using LLMs for evaluation, prompt improvement, and large-scale structured metadata generation, which is directly relevant to builder workflows around AI systems and automation. It is especially useful because it shows an applied loop where evaluation outputs feed prompt optimization and data creation rather than treating LLMs as a one-shot generator. Concepts: LLM Evals Context Engineering Entities: DoorDash Spark Source: careersatdoordash.com
Knowledge Gaps
Topics the AI stream keeps raising that the knowledge base hasn’t sufficiently covered yet — candidates for what to learn next. Agent Memory · Context Engineering · RAG · Embeddings · Reranking
🚀 Models & Releases (4)
prio 7Bilibili releases Index-1.9B open small language model series Concepts: Open Source LLMs RAG LLM Evals Entities: Bilibili Index-1.9B Index-1.9B-Base Index-1.9B-Pure Source: arxiv.orgprio 7Sber open-sources GigaAM Multilingual and GigaChat Audio Concepts: Long Context LLM Evals Entities: Sber GigaAM Multilingual GigaChat Audio GigaChat-Max-Audio Source: interspeech2026.orgprio 6Alice AI ART 2.0 moves toward a unified image generation and editing model Concepts: LLM Evals Entities: Alice AI Habr Shadewrum Alice AI ART 1.0 Source: habr.comprio 6Bonsai 27B: 1-Bit Quantization Enables 27B-Class Agents on Phones Concepts: Agents Tool Use LLM Evals Open Source LLMs Entities: Bonsai 27B Qwen3.6-27B Bonsai 8B Source: prismml.com
🧪 Research Papers (136)
prio 11Valid Does Not Mean Necessary: A Paper on CoT Inefficiency Concepts: LLM Evals Source: arxiv.orgprio 11QIMG-7 benchmark exposes how polluted multimodal RAG can fail, and SATR adds source-aware trust selection Concepts: RAG RAG Evaluation LLM Evals Entities: gpt-4o-mini Source: arxiv.orgprio 11Benchmarking faithfulness in LLM-generated clinical trial summaries Concepts: LLM Evals RAG RAG Evaluation Entities: OpenAI Anthropic Google ClinicalTrials.gov Source: arxiv.orgprio 11What Coding Agents Need in Context to Edit Code Concepts: Code Agents Context Engineering Long Context LLM Evals Source: arxiv.orgprio 11Eval-Pair Matrix studies same-model bias in grounded RAG judging Concepts: RAG LLM Evals RAG Evaluation Entities: GPT Grok Gemini Source: arxiv.orgprio 10Coresets Before Score Sets: Selecting Prompt Subsets for LLM Benchmarks Concepts: LLM Evals Embeddings Source: arxiv.orgprio 10Compile First: Executable SOP Programs and Capability-Gated Runtime for LLM Agents Concepts: Agents Context Engineering LLM Evals Source: arxiv.orgprio 10Reference-Based Distillation Detection in LLMs Concepts: LLM Evals Entities: OpenAI DeepSeek QwQ DeepSeek R1 Source: arxiv.orgprio 10Associative recurrent memory for extending LLM context Concepts: Long Context Entities: Associative Recurrent Memory Transformer 2 sources: arxiv.org, arxiv.orgprio 10PaperRouter-Agent uses folder members to route new papers in personal libraries Concepts: Agents RAG Source: arxiv.orgprio 10MemDecay proposes region-aware KV cache eviction for LLM agents Concepts: Agents Context Engineering Entities: Qwen2.5-1.5B Qwen2.5-3B Source: arxiv.orgprio 10ToFu: A White-Box, Token-Efficient Agent Harness for Researchers Concepts: Agents Code Agents Tool Use Context Engineering Source: arxiv.orgprio 10Benchmarking LLM Legal Citation Fabrication in GDPR and Saudi PDPL Concepts: LLM Evals Entities: arXiv Gemini 2.5 Flash GPT-OSS 120B Nemotron-3-Super-120B Source: arxiv.orgprio 10ResearchQA benchmarks citation-grounded QA on scientific papers Concepts: LLM Evals RAG Evaluation Source: arxiv.orgprio 10LLM Jury Consensus Can Beat Self-Consistency for Reasoning Selection Concepts: LLM Evals Entities: arXiv alphaXiv Source: arxiv.orgprio 10PASB benchmarks persistent sycophancy in stateful personal agents Concepts: Agent Memory Agents LLM Evals Source: arxiv.orgprio 10SPARK profiles latent reasoning states and steers under-activated examples Concepts: LLM Evals Entities: Qwen3 series Qwen3-4B Qwen3-8B Source: arxiv.orgprio 10Survey maps metacognition in LLMs Concepts: LLM Evals Entities: DAIR.AI arXiv Source: arxiv.orgprio 10Structured Thoughts adds try/outcome reasoning blocks and context pruning Concepts: Context Engineering Source: arxiv.orgprio 10FATE is a specialized 8B model for evaluating AI tutors Concepts: LLM Evals Entities: arXiv FATE Gemini 2.5 Flash ChatGPT 5.5 Instant Source: arxiv.orgprio 10Legal information processing paper reports gains from ensembles, reranking, and retrieval-augmented prompting Concepts: RAG Reranking Hybrid Search Entities: Team DU arXiv COLIEE Qwen3-235B 2 sources: arxiv.org, arxiv.orgprio 10Quantization can change LLM reasoning failure modes without hurting accuracy Concepts: LLM Evals Entities: Llama-3.2-3B Source: arxiv.orgprio 10A Theory of Least Autonomy for Agentic AI Systems Concepts: Agents Tool Use Source: arxiv.orgprio 10Abstention-aware RL for search agents Concepts: RAG LLM Evals Tool Use RAG Evaluation Source: arxiv.orgprio 10EYT-Bench benchmarks multi-turn dialogue with decoupled simulation, modeling, and judging Concepts: LLM Evals Entities: Gemma 4 GPT 5.5 DeepSeek V4 Pro Source: arxiv.orgprio 10EvidentialRAG adds uncertainty-aware conflict handling to multi-source RAG Concepts: RAG RAG Evaluation Source: arxiv.orgprio 10Probing LLM internal states to detect confident hallucinations in financial QA Concepts: LLM Evals Entities: Qwen3-8B Llama-3.1-8B Gemma-2-9B Source: arxiv.orgprio 10Google DeepMind on evaluating model routers beyond accuracy and cost Concepts: LLM Evals Entities: DAIR.AI Google DeepMind Source: arxiv.orgprio 9RAGU proposes a multi-step GraphRAG pipeline with a compact extractor model Concepts: RAG LLM Evals Entities: MIT Meno-Lite-0.1 Qwen2.5 32B Source: arxiv.orgprio 9A framework for scaling medical agents from assistance to autonomy Concepts: Agents Tool Use LLM Evals Source: arxiv.orgprio 9Paper examines how sampling temperature affects ideological discourse transfer in RAG Concepts: RAG RAG Evaluation LLM Evals Entities: arXiv 2 sources: arxiv.org, arxiv.orgprio 9Benchmarking a Teaching-Feedback Classification Protocol Across Embeddings and Languages Concepts: Embeddings LLM Evals Entities: arXiv Source: arxiv.orgprio 9AI YOU updates persona profiles with Bayesian inference, conformal prediction, and layered memory Concepts: Agent Memory Agents LLM Evals Source: arxiv.orgprio 9MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment Concepts: LLM Evals Source: arxiv.orgprio 9ABot-AgentOS proposes a robotic agent OS with persistent multi-modal memory Concepts: Agents Agent Memory Tool Use LLM Evals Source: arxiv.orgprio 9CRiT-QA adds counterfactual traps for multi-hop reasoning eval Concepts: LLM Evals Source: arxiv.orgprio 9Paper separates value disagreement from response determinism in cross-model comparisons Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 9Study Finds Frontier Models Struggle With Information Seeking in Agentic Clinical Reasoning Concepts: Agents LLM Evals Source: arxiv.orgprio 9Survey maps how agent skill libraries evolve over time Concepts: Agents Tool Use Source: arxiv.orgprio 9QwenPaw-Data proposes a three-part architecture for enterprise data agents Concepts: Agents RAG Entities: arXiv Source: arxiv.orgprio 9Demographic Prompting Study Finds Too Many Attributes Can Hurt LLM-Human Agreement Concepts: Context Engineering Source: arxiv.orgprio 9Opti-Agent-Bench benchmarks end-to-end optimization R&D agents on business problems Concepts: LLM Evals Agents Code Agents Source: arxiv.orgprio 9AgentFootprint evaluates storage overhead in LLM agent runs Concepts: LLM Evals Agents Source: arxiv.orgprio 9Exploring Agentic Workflows for Generating High Quality Math Visual Aids Concepts: Agents LLM Evals Source: arxiv.orgprio 9Verification of Adaptive Agentic Controllers through Finite Rule Revision Concepts: Agents LLM Evals Source: arxiv.orgprio 9Hallucination Detection with Diversion Decoding Concepts: LLM Evals Source: arxiv.orgprio 9GEIS proposes a generation-evaluation-improvement loop for long-form article generation Concepts: Agents Tool Use LLM Evals Context Engineering Entities: Tasi Harness STORM Source: arxiv.orgprio 9Constraint-Aware Hierarchical Search for regulation-driven fine-grained classification Concepts: LLM Evals Entities: arXiv arXivLabs alphaXiv CatalyzeX Source: arxiv.orgprio 9LightMem-Ego: A Lightweight Multimodal Memory System for Everyday Life Assistants Concepts: Agent Memory Source: arxiv.orgprio 9Kahneman4Review benchmarks LLM judge reliability on peer reviews Concepts: LLM Evals Source: arxiv.orgprio 9Matched Evaluation Finds SAE Safety Control Is Regime-Dependent Concepts: LLM Evals Entities: gemma-2-9b-it Source: arxiv.orgprio 9Paper argues that sign-branched repetition penalties depend on logit centering and can break structured output Concepts: LLM Evals Entities: Hugging Face Source: arxiv.orgprio 9Study Compares LLM-Generated EFL Grammar Drill Formats and CEFR-J Difficulty Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 8Microsoft rollout study finds peer-driven adoption of Claude Code and GitHub Copilot CLI Concepts: Code Agents Entities: Microsoft Anthropic GitHub Source: arxiv.orgprio 8Calibrated e-CUSUM decoding for quantized reasoning models Concepts: LLM Evals Entities: DeepSeek-R1-Distill-Qwen-1.5B Source: arxiv.orgprio 8HardChoices probes whether LLMs can take coherent stances on societal issues Concepts: LLM Evals Source: arxiv.orgprio 8AdvancedMathBench benchmarks advanced mathematical proof generation and verification Concepts: LLM Evals Entities: GPT-5.5 xhigh Source: arxiv.orgprio 8Paper argues LLM planning is not one skill, but two Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 8Continual fact writing into model weights appears fragile compared with context Concepts: Context Engineering Entities: arXiv Qwen3 Source: arxiv.orgprio 8The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese Concepts: LLM Evals Source: arxiv.orgprio 8The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory Concepts: Agent Memory Agents LLM Evals Source: arxiv.orgprio 8Improved Answer Selection with Pre-Trained Word Embeddings Concepts: Embeddings Source: arxiv.orgprio 8EAST benchmarks theory-of-mind failures in LLM dialogue Concepts: LLM Evals Source: arxiv.orgprio 8Modular framework for evaluating LLM alignment in mental health Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 8LOGOS proposes a governance layer for evolving AI agent teams Concepts: Agents Tool Use LLM Evals Source: arxiv.orgprio 8Harness-Native Agentic Routing Turns Execution Traces Into Training Data Concepts: Agents Tool Use Context Engineering Entities: arXiv.org OpenSquilla LightGBM Source: arxiv.orgprio 8Length penalties reduce chain-of-thought visibility without removing hint influence Concepts: LLM Evals Entities: Qwen3-4B Qwen3-14B Source: arxiv.orgprio 8AuditWeave: tamper-evident evidence logging for AI-assisted workflows Concepts: RAG Source: arxiv.orgprio 8Metadata-Free Meta-Reweighted DPO for Noisy Preference Labels Entities: Anthropic arXiv Source: arxiv.orgprio 8EvoCUA-1.5: Online RL for Multi-turn Computer-Use Agents Concepts: Agents Source: arxiv.orgprio 8Single-pass Unity code generation study finds no compilable scene across 10,400 trials Concepts: LLM Evals Source: arxiv.orgprio 8BatteryLake turns heterogeneous battery aging data into benchmark-ready assets Concepts: Agents LLM Evals Source: arxiv.orgprio 8Faithful, Not Corrective: Message Format Effects in Multi-Hop Agent Relays Are Tier-Dependent Concepts: Agents Context Engineering Source: arxiv.orgprio 8Claude Fable 5 evaluation on biomedical benchmarks highlights refusal behavior Concepts: LLM Evals Entities: Anthropic Claude Fable 5 GPT-5 Source: arxiv.orgprio 8Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory Concepts: Agents Tool Use Agent Memory Source: arxiv.orgprio 8STAMP adds provenance-based credit assignment for deep search agents Concepts: Agents LLM Evals Source: arxiv.orgprio 8Agentic context learning appears to depend on specification acquisition, not just retrieval Concepts: Context Engineering LLM Evals Long Context Entities: GPT-5.1 Qwen3.5-27B Gemini 3 Pro Source: arxiv.orgprio 8Imaging-101 benchmarks LLM coding agents on computational imaging Concepts: Code Agents LLM Evals Source: arxiv.orgprio 8Norm Enforcement for AI Agents studies how enforcement mechanisms shape behavior in multi-agent systems Concepts: Agents Source: arxiv.orgprio 7NextFund Introduces a Live Evaluation Platform for Agentic Portfolio Management Concepts: Agents LLM Evals Source: arxiv.orgprio 7OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping Concepts: LLM Evals Entities: arXiv arXiv.org arXivLabs alphaXiv Source: arxiv.orgprio 7modelDNA verifies open-weight model lineage from sampled weight fingerprints Concepts: Open Source LLMs Entities: Hugging Face Source: arxiv.orgprio 7M+Adam proposes additive-multiplicative optimization for low-precision training Entities: LLaMA Source: arxiv.orgprio 7SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL Concepts: Agents Tool Use Context Engineering LLM Evals Source: arxiv.orgprio 7STEC compresses evidence for final answer selection in multi-hop QA Concepts: RAG Agents Source: arxiv.orgprio 7FlashTrie speeds up constrained beam search for generative retrieval on GPUs Concepts: RAG Source: arxiv.orgprio 7UNIBROWSE proposes a data pipeline for multimodal BrowseComp agents Concepts: Agents Tool Use LLM Evals Entities: Qwen3.5-35B-A3B GPT-5 Gemini 2.5 Pro Gemini 2.5 Flash Source: arxiv.orgprio 7Pitfalls of Administrative Censoring in Survival Models with Time-Indexed Inputs Source: arxiv.orgprio 7Automated Textbook Auditing with Multi-Agent LLM Systems Concepts: Agents Source: arxiv.orgprio 7Agentic-DPO turns expert trajectories into state-level preference training for agents Concepts: Agents Tool Use LLM Evals Source: arxiv.orgprio 7Can Agentic Trading Systems Pay for Their Own Intelligence? Concepts: Agents LLM Evals Entities: DeepSeek-V3.2 GLM-4.7 Source: arxiv.orgprio 7Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders Concepts: LLM Evals Entities: DeepSeek-V3.1 Judgmental Llama Judgmental OLMo Source: arxiv.orgprio 7LLM Opinion Summarization with Stratified Sampling and Token Efficiency Entities: Amazon Tripadvisor X/Twitter Source: arxiv.orgprio 7On-device subtitle translation optimized around quantization and vocabulary size Entities: Google Apple LMT-60-0.6B GPT-4o Source: arxiv.orgprio 7MawForge studies bounded local inference for Mixture-of-Experts models Source: arxiv.orgprio 7RouteCast evaluates model-generated strategic routes when ground truth arrives late Concepts: LLM Evals Source: arxiv.orgprio 7CLIR-Bench benchmarks multimodal QA over irregular clinical time series Concepts: LLM Evals Source: arxiv.orgprio 7MET proposes theory-grounded multilingual moral reasoning with a new benchmark Concepts: LLM Evals Entities: Qwen3-4B Qwen3-8B Gemma3-4B Source: arxiv.orgprio 6Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Entities: Google alphaXiv 2 sources: twitter.com, alphaxiv.orgprio 6Budgeted placement of strong correctors in weak multi-agent swarms Concepts: Agents LLM Evals Entities: Qwen3 Source: arxiv.orgprio 6AutoVSR uses VLMs plus symbolic solving to derive circuit expressions from schematics Concepts: Agents Tool Use Entities: VLMs Source: arxiv.orgprio 6Weight-Adjusted Gradients for Parameter Importance in LLMs Source: arxiv.orgprio 6Token probability differences between production and perception prompts in LLMs Entities: arXiv Llama Llama-3.1-8B EuroLLM-9B Source: arxiv.orgprio 6NVAITC AI Scientist proposes a governed agentic research system for biomedical workflows Concepts: Agents Tool Use Entities: NVAITC Source: arxiv.orgprio 6SDABench evaluates LLMs on scientific data analysis across six capability types Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 6ARMOR proposes off-policy anchor samples to stabilize on-policy LLM RL Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 6JobHop v2 release: a large-scale career trajectory dataset extracted from multilingual resumes Concepts: LLM Evals Entities: VDAB ESCO Source: arxiv.orgprio 6On the modality gap in CLIP-style contrastive learning Concepts: Embeddings LLM Evals Entities: CLIP Source: arxiv.orgprio 6Large language model agents for inverse design of MOFs for gas separation Concepts: Agents Agent Memory Source: arxiv.orgprio 6Cross-Layer Misalignment Detection for Agent Skills Concepts: Agents LLM Evals Source: arxiv.orgprio 6RDQ proposes cascaded error compensation for low-bit LLM quantization Entities: Qualcomm arXiv LLaMA-3-8B Qwen 2.5 7B Source: arxiv.orgprio 6LLMs struggle with cooperative communication under partial information Source: arxiv.orgprio 6EasyOPD presents an on-policy distillation framework for LLMs Entities: arXiv Source: arxiv.orgprio 6Governed answer reuse via quotient classes instead of embedding similarity Concepts: Embeddings Source: arxiv.orgprio 6Relational Positioning Measures Multi-Turn Dialogue Risk Concepts: LLM Evals Source: arxiv.orgprio 6UMoE proposes expert pruning and regrowth before domain fine-tuning Entities: Qwen Qwen3-30B-A3B Qwen3.5-35B-A3B Qwen3-30B-A3B-Thinking Source: arxiv.orgprio 6Explicit Reasoning Can Hurt Patent Claim Drafting Concepts: LLM Evals Source: arxiv.orgprio 6Diagnosing Thinking Collapse in On-Policy Self-Distillation Entities: arXiv alphaXiv Connected Papers Litmaps Source: arxiv.orgprio 6ProgramTab uses Python preprocessing and SQL extraction for table reasoning Entities: arXiv Source: arxiv.orgprio 6Cost of Reasoning in non-English Languages: Japanese reasoning study with Qwen-3-Swallow-8B Concepts: LLM Evals Entities: Qwen-3-Swallow-8B Qwen-3-8B Source: arxiv.orgprio 6Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems Concepts: Agents Entities: arXiv alphaXiv CatalyzeX DagsHub Source: arxiv.orgprio 6Architecture-dependent reasoning scaffolds in Hotelling spatial markets Concepts: LLM Evals Entities: GPT-4.1-mini GPT-5-mini Source: arxiv.orgprio 6Global Merger-Arbitrage Forecasting with Language Models Concepts: Context Engineering Long Context LLM Evals Source: arxiv.orgprio 6PolyInterview: An LLM Platform for Mock Interviews with Multimodal Assessment Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 6RUBRIC: Realism-Utility Balanced Ranking for Imbalanced Classification Source: arxiv.orgprio 6TreeThink: Asynchronous Tree Search for LLM-Based Theorem Proving Source: arxiv.orgprio 6A paper proposes new abstractiveness metrics for summarization evaluation Concepts: LLM Evals Entities: BART-large-cnn Pegasus-xsum DistilBart MT5-small Source: arxiv.orgprio 6Paper separates preprocessing instability from measurement instability in LLM-based stance analysis Concepts: LLM Evals Entities: YouTube arXiv Source: arxiv.orgprio 6New Slavic-language corpus for persuasion-technique detection Concepts: LLM Evals Source: arxiv.orgprio 6Similarity-based source-language selection for low-resource ASR on Warlpiri Entities: Whisper Source: arxiv.orgprio 6Predicting When Low-Precision Training Freezes Entities: GPT GPT-2 Source: arxiv.orgprio 6Query-Focused Event Summarization introduces a new dataset and benchmark Concepts: RAG Entities: arXiv Source: arxiv.orgprio 6The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions Entities: SmBoP RASAT Source: arxiv.orgprio 6Position paper argues ground truth is constructed, not objective Concepts: LLM Evals Entities: arXiv Source: arxiv.orgprio 6Output-Aware Safety Guardrails for Multimodal LLMs Source: arxiv.orgprio 6Cobras frames activation steering as a Schrödinger Bridge Source: arxiv.org
🛠 Tools & Frameworks (15)
prio 9Picchio benchmarks local LLMs with effective bits-per-weight, three speed lanes, and GPU fallback detection Concepts: Open Source LLMs Entities: Hugging Face LM Studio NVIDIA Qwen3.5-9B Source: github.comprio 8Claude Code plugin plays audio cues when it is waiting on you Entities: Claude Code Hacker News Source: github.comprio 8A cache-friendly uvx pattern for GitHub Actions Entities: GitHub PyPI Source: simonwillison.netprio 8KOMPAS Guard uses a 34M bilingual encoder to map task descriptions to KOMPAS API elements Concepts: Agents Tool Use Entities: USER2-small Source: habr.comprio 8Codex CLI regression hides human-readable subagent task text after message encryption Concepts: Agents Tool Use Source: github.comprio 8Customizing Claude message display with a hook Concepts: Tool Use Entities: Bluesky Hacker News Source: jola.devprio 8Agnost AI launches to surface failures from real agent conversations Concepts: Agents Tool Use LLM Evals Entities: Agnost AI Y Combinator MCP Toolbox for Databases OpenTelemetry Source: agnost.aiprio 7Jacquard 0.1: a language for reviewing AI-written code Entities: FriendMachine Source: github.comprio 7CLI that turns YouTube guitar lessons into PDF tabs Entities: Anthropic YouTube Claude Sonnet 5 Source: github.comprio 7Kadr: an open-source video editor built around Claude Code and MCP Concepts: Agents Code Agents MCP Tool Use Source: habr.comprio 7Cursor IDE: Unpatched Remote Code Execution Vulnerability Via Malicious git.exe Entities: Cursor Mindgard HackerOne Source: mindgard.aiprio 6DOOMQL: a Doom-like game driven by SQL in SQLite Entities: OpenAI Anthropic GitHub GPT-5.6 Sol Source: simonwillison.netprio 6JetBrains introduces YouTrackDB as a general-purpose object-oriented graph database Entities: JetBrains Source: github.comprio 6Microsoft Entra ID to make passkeys the default authentication method starting in September Entities: Microsoft BleepingComputer Source: bleepingcomputer.comprio 6alphaXiv highlights a runnable Autoresearch workflow for a paper on memorized knowledge and finetuning Concepts: Agents Tool Use Entities: alphaXiv Source: alphaxiv.org
🏢 Industry / Business (2)
prio 6LimX Dynamics raises $200M Pre-IPO round and frames its robot stack around brain-skill separation Concepts: Agents Agent Memory Entities: 逐际动力 IDG资本 蓝思科技 GGG Group Source: qbitai.comprio 6Phishing kits target Microsoft 365 accounts by bypassing MFA Entities: BleepingComputer Microsoft ReliaQuest SharePoint Source: bleepingcomputer.com
💬 Opinions (22)
prio 12Building a Claude Code-like agent in Python and the engineering choices that make it usable Concepts: Code Agents Tool Use Context Engineering Entities: Claude Code Claude API Aider Amp 3 sources: habr.com, habr.com, habr.comprio 12Claude Code Cost Breakdown Shows Cached Context Dominates Spend Concepts: Context Engineering Entities: Opus Fable Source: habr.comprio 11Building an AI Agent Team for Real Business, Not Just Chat Demos Concepts: Agents Agent Memory Tool Use Context Engineering Source: habr.comprio 10A CCA-F mock exam exposed a common mistake: treating prompts as architecture Concepts: LLM Evals Entities: Anthropic Source: habr.comprio 10Practical news personalization without big data: ranking Telegram with pgvector and five signals Concepts: Embeddings Vector Database Entities: BAAI/bge-m3 Source: habr.comprio 9The case for replacing code review with system checks in AI-assisted development Concepts: Code Agents LLM Evals Entities: AURAIDO Source: habr.comprio 9Eight Practical Lessons from Running Open LLMs in Production Concepts: Open Source LLMs Entities: WB-Tech Ollama llama.cpp Source: habr.comprio 9A manifesto for working with AI coding agents Concepts: Code Agents Source: habr.comprio 9Building a Memory-Enhanced Telegram Bot with RAG: A Terminator Chatbot Story Concepts: Agents Agent Memory RAG Entities: OpenRouter NVIDIA Nemotron-3-Super-120B Source: habr.comprio 8Agents.md: inspect first, verify before changing code Concepts: Code Agents Source: gist.github.comprio 8The Definition of Done for Software Work Source: adi.bioprio 8Why AI-Accelerated Development Can Increase Production Failures Entities: Habr Source: habr.comprio 7Measuring AI referral traffic in Yandex Metrika and why ChatGPT with Alice do not show up Entities: Yandex OpenAI Perplexity Google Source: habr.comprio 7Git history commands aim to simplify rewriting commit stacks Source: lalitm.comprio 7How a gas-station service audit system combines archived video and audio Entities: GigaAM v3 Source: habr.comprio 7How I Used AI to Build an OVAL Viewer Concepts: Code Agents Entities: OpenSCAP Source: habr.comprio 7From requirements to diagnosis: how the AI delivery model is changing Source: habr.comprio 7Show HN: RL-trained agent that trains models with RL Concepts: Agents Tool Use LLM Evals Entities: Tinker Runpod Hugging Face Qwen3.6-35B-A3B Source: github.comprio 6Benchmarking YOLO inference on Raspberry Pi 5 HAT+ with HAILO-8L Entities: Xiaomi Hailo YOLO8n YOLO8s Source: habr.comprio 6ICML 2026 keynote argues for AI adaptation, not a sudden job cliff Concepts: LLM Evals Agents Source: normaltech.aiprio 6Tensor Is the Might Entities: GPT-5 Source: zserge.comprio 6AI Agents and the Tower of Babel: How Removing Friction Erodes Shared Understanding Concepts: Agents Code Agents Source: lucumr.pocoo.org
📦 Other (1)
prio 7Habr guide on writing image-generation prompts Entities: Хабр Source: habr.com
FAQ
What is in the 2026-07-14 AI brief?
The 2026-07-14 brief selected 185 signal items for AI builders and filtered 270 items as noise, using the radar’s community-relevance scoring.