Skip to content

Type: open-access research repository

arXiv is an open-access repository hosting preprints across physics, mathematics, computer science, and quantitative biology, operated by Cornell University. Much frontier AI research circulates on arXiv before formal peer review. GROUNDING tracks notable arXiv papers, AI methods, and the arXivLabs tools that surface code and discussion around them.

Recent Updates

  • 2026-07-24: AsymVerify uses confidence-gated verification to detect political evasion (cs.CL updates on arXiv.org) · github.comLLM Evals SemEval-2026 CLARITY ACL Kaons Sebastien Kawada · GLM-4.7
  • 2026-07-24: ReliableTableQA studies how much supervision reliability annotation needs (cs.LG updates on arXiv.org) · arxiv.orgLLM Evals H&M
  • 2026-07-24: Paper studies how lie type, depth, and sparsity affect deception probes in LLMs (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals
  • 2026-07-24: Training-Free Routing for Local-Cloud LLM Collaboration (cs.AI updates on arXiv.org) · arxiv.org
  • 2026-07-24: PersonaTrail benchmarks personalized web agents with browsing-history memory (cs.AI updates on arXiv.org) · arxiv.orgAgent Memory Agents
  • 2026-07-24: Polynomial-time control for LLMs under LR(k) grammar constraints (cs.AI updates on arXiv.org) · arxiv.org — Maximilian Scribner
  • 2026-07-24: Evaluating Whether LLMs Can Detect Their Own Generated Content (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals
  • 2026-07-24: Paper measures ideological drift in news-grounded LLM QA (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals RAG Evaluation QBias
  • 2026-07-24: Autonomous Topology Mutation for multi-agent LLM systems (cs.AI updates on arXiv.org) · arxiv.orgAgents Tool Use alphaXiv · Connected Papers · Litmaps · scite · DagsHub · Gotit.pub · Hugging Face · ScienceCast · CORE · arXivLabs · DeepSeek-V3
  • 2026-07-24: EvoSQL adds memory-guided multi-round search for Text-to-SQL (cs.AI updates on arXiv.org) · github.comAgents Tool Use Hugging Face · Qwen3-4B · Qwen2.5-Coder-3B
  • 2026-07-24: Inference-time knowledge injection improves zero-shot delirium prediction in open-weight LLMs (cs.CL updates on arXiv.org) · arxiv.orgContext Engineering LLM Evals Llama-3.1-8B · Llama-3.3-70b · GPT-5.2
  • 2026-07-24: Learn2Zinc fine-tunes small language models for MiniZinc text-to-model translation (cs.CL updates on arXiv.org) · arxiv.orgQwen3 · LLaMA · Gemma · GPT-OSS
  • 2026-07-24: Monkey King Bang: A Unified Scientific Multimodal Foundation Model (cs.LG updates on arXiv.org) · github.comHugging Face · Meta sais-org MKB · Qwen3-VL · Qwen3-VL-8B ESM-2 ConvFormers Swin-ViT · SAM 3 Biology-Instructions · Llama-3.1-8B Intern-S1-Pro BiomedParse HRES
  • 2026-07-24: ArXiv study compares regex filtering and LLM alignment under adversarial probes (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals Google Alexandre Maiorano · Gemini 2.5 Flash
  • 2026-07-24: Workload-aware caching for multi-agent pipelines (cs.AI updates on arXiv.org) · arxiv.orgAgents
  • 2026-07-24: ExecuGraph evaluates a multi-agent, execution-checked backend code synthesis workflow (cs.AI updates on arXiv.org) · arxiv.orgAgents Code Agents Tool Use LLM Evals RAG Ollama DeepSeekCoder V2 Lite
  • 2026-07-24: FlowEdit proposes information-theoretic control for branch-aware LLM reasoning (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals alphaXiv · Connected Papers · Litmaps · scite · DagsHub · Gotit.pub · Hugging Face · ScienceCast · CORE
  • 2026-07-24: RL-trained vision-language critic for UI quality violations (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals
  • 2026-07-24: Paper studies personality steering in LLMs with Jungian cognitive functions (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals Llama-3.1-8B
  • 2026-07-24: CANN Bench benchmarks agent-generated kernels on Huawei Ascend NPUs (cs.AI updates on arXiv.org) · arxiv.orgCode Agents LLM Evals Huawei
  • 2026-07-24: REFACT proposes adaptive fact restatement for grounded reasoning (cs.CL updates on arXiv.org) · github.comLLM Evals Long Context GitHub
  • 2026-07-24: CSPF: Constrained Fusion for Non-Verifiable Preference Evaluation (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals
  • 2026-07-24: LegalCiteTrust benchmarks citation trustworthiness in Chinese legal research reports (cs.CL updates on arXiv.org) · arxiv.orgLLM Evals RAG Evaluation Agents
  • 2026-07-24: ConfidenceBench evaluates verbalized confidence calibration in 15 LLMs (cs.AI updates on arXiv.org) · arxiv.orgLLM Evals Claude Opus 4.6 · Gemini 3.1 Pro Preview Gemini 3.1 Flash-Lite
  • 2026-07-24: From Word-Level Dictionary to Sentence-Level Semantics for Multilingual Grievance Labeling (cs.CL updates on arXiv.org) · github.comLLM Evals Hugging Face

FAQ

What is arXiv?

arXiv is an open-access repository hosting preprints across physics, mathematics, computer science, and quantitative biology, operated by Cornell University. Much frontier AI research circulates on arXiv before formal peer review. GROUNDING tracks notable arXiv papers, AI methods, and the arXivLabs tools that surface code and discussion around them.

What does this page track?

Dated radar mentions, source links, related concepts, and builder-relevant context for arXiv, collected automatically by GROUNDING.

When was arXiv last mentioned?

arXiv was most recently mentioned in a radar update dated 2026-07-24.