Skip to content

Type: technology company

Alibaba is a Chinese multinational technology conglomerate whose cloud division develops the open-weight Qwen model family and offers AI infrastructure through Alibaba Cloud. Its open releases are influential across the global open-model ecosystem. GROUNDING tracks Alibaba’s Qwen models, cloud AI services, and open-weight strategy.

Recent Updates

  • 2026-08-18: Six LLM Judges from Different Labs Show Only 1.9 Independent Voices Due to High Error Correlation (Все статьи подряд / Искусственный интеллект / Хабр) · habr.comLLM Evals RAG Evaluation DeepSeek · Meta · Mistral · Amazon Ladha Boland · DeepSeek-Chat Qwen-2.5–7B Qwen-2.5–72B Llama-3.3–70B Mistral Large Amazon Nova-Pro
  • 2026-08-20: Wall Street Benchmark: Alibaba Qwen Office Agents Lead in Real-World Evaluation; Engineering Quality Rivals Model Capability (量子位) · qbitai.comAgents Tool Use Context Engineering Jefferies Anthropic · OpenAI · Qwen · Qwen 3.8 Max
  • 2026-08-20: ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (cs.AI updates on arXiv.org) · arxiv.orgAgents LLM Evals OpenAI · Google · GPT-5.4 · GPT-5.4-mini · GPT-5-mini · Gemini 3 Flash · Gemini 3.1 Flash-Lite Qwen3-VL-235B UI-TARS-1.5-7B
  • 2026-08-20: Cache-Aware Routing for Large Mixture-of-Experts Models on Edge Hardware: A Pre-Registered Study of the Quality-Efficiency Tradeoff (cs.AI updates on arXiv.org) · arxiv.orgOpen Source LLMs Shriniwas Ramesh Suram Qwen3-235B · Qwen3 30B
  • 2026-08-20: Persona-Guided LLM Agents for Task-Oriented Dialogue (cs.CL updates on arXiv.org) · arxiv.orgAgents OpenAI · Google Maryam Shoaeinaeini · GPT-4o Qwen3-Next-80B · Gemini 2.0 Flash
  • 2026-08-20: Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks (Hacker News) · dreadnode.ioLLM Evals Anthropic · OpenAI · Google · xAI · DeepSeek · z.ai · NIST HackTheBox · e2b Michael Kouremetis · Claude Opus 4.8 · Claude Opus 4.7 · Claude Opus 4.6 · Claude Sonnet 5 · Claude Sonnet 4.6 · Claude Haiku 4.5 · GPT 5.5 · GPT-5.4 · GPT-5.4-mini · Gemini-3.1-Pro · Gemini 3 Flash · Grok 4.20 · Grok 4.3 · DeepSeek V4 Pro · DeepSeek-R1-0528 · DeepSeek-V4-Flash Qwen 3-7 Max Qwen 3.6 Max · Qwen 3.6 Plus Qwen3-Coder-Next · GLM-5.1 GLM-5-turbo
  • 2026-08-22: Frontier Model Autonomous Agent Benchmark: NanoGPT Optimizer Speedrun Leaderboard (breakingnewsofficial) · primeintellect.aiAgents Code Agents LLM Evals Anthropic · OpenAI · Moonshot AI · xAI · Zhipu · DeepSeek · Muse · Fable 5 · Opus 5 · Kimi K3 · Opus 4.8 · GPT-5.6 Sol GPT-5.6 Sol Pro · Sonnet 5 · GPT-5.6 Luna · Grok 4.5 · Qwen 3.8 Max · GLM-5.2 · DeepSeek V4 Pro · GPT-5.6 Terra · Grok 4.6 · Muse Spark 1.2 · Muse Spark 1.1 · GPT 5.5 · Kimi K2.7
  • 2026-08-23: Rooting an Amazon Fire tablet with four frontier LLMs: how Kimi K3 found CVE-2022-38181 (breakingnewsofficial) · ericpardee.github.ioAnthropic · Moonshot AI · Amazon · ARM · GitHub Man Yue Mo · Claude · Kimi K3 · GLM-5.2 · GLM-5.3
  • 2026-08-23: Fable’s High Cost Drives Shift to Optimized Workflows Over Premium Models (breakingnewsofficial) · dbreunig.comAnthropic · Fable · Opus · GLM-5.2 K3 · Qwen 5.6
  • 2026-08-25: Qwen3.8-27B: Open-Weight Model Outperforms Claude on Legal Benchmarks and Emerges as Capable Agent Model (breakingnewsofficial) · orcarouter.aiAgents RAG Open Source LLMs Anthropic Dario Amodei · Qwen3.8-27B · Claude Opus 4.8
  • 2026-08-25: Qwen3.8-27B Achieves Top Ranking in Image-to-Code Generation (breakingnewsofficial) · t.meOpen Source LLMs Qwen3.8-27B · Kimi K3
  • 2026-08-26: Qwen3.8-Flash: Open-Weight Multimodal MoE Model with Competitive API Pricing (breakingnewsofficial) · x.comOpen Source LLMs Qwen Qwen3.8-Flash Qwen4
  • 2026-08-28: AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling (breakingnewsofficial) · arxiv.orgLLM Evals Agents Tool Use OpenAI · Google · GPT-5.4 · Gemini 2.5 Pro QwQ-32B · GPT-OSS 120B
  • 2026-08-29: Alibaba Qoder launches desktop agent with voice interaction and autonomous code verification (breakingnewsofficial) · qbitai.comCode Agents Agents Context Engineering Codebase Indexing Tool Use Qwen 3.8
  • 2026-09-01: E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operations (breakingnewsofficial) · arxiv.orgAgents LLM Evals OpenAI · Anthropic · Zhipu · DAIR.AI · GPT-5.6 Sol · Fable 5 Qwen3.8-Max-Preview · GLM-5.2
  • 2026-09-02: Alibaba Updates Qwen3.8-Max Flagship Model, Frontend Programming Ability Ranks First Globally (breakingnewsofficial) · qbitai.com — Qianwen Qwen3.8-Max Claude Opus5 · Kimi K3
  • 2026-09-04: What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation (breakingnewsofficial) · arxiv.orgLLM Evals Context Engineering OpenAI · gpt-oss-20b · GPT-OSS 120B · GPT-5.4-mini · Qwen3.5-9B · Qwen3.5-27B Qwen3.5-122B
  • 2026-09-07: Hidden state bridge enables 4B mobile model to match larger cloud model on ARC-AGI 3 (breakingnewsofficial) · qbitai.comAgents Context Engineering Mostik Zhipu · Anthropic · OpenAI Stanislav Smirnov Sasha Malysheva · Qwen-3.5 · Qwen 3 · GLM-5.2 · GPT-6 Astra Claude-3.5-Haiku
  • 2026-09-07: When Frontier LLMs Act as Autonomous Agents: Fraud, Spam, and Rationalization (breakingnewsofficial) · bottlenecklabs.comAgents Tool Use Stripe · Exa · Browserbase Playwriter Inkbox Mailjet · Hacker News · GitHub Meow.com · OpenCode · Qwen 3.8 · Grok 4.5 · Muse
  • 2026-09-08: TradingAgents: Open-source multi-agent framework for AI-powered financial analysis (breakingnewsofficial) · github.comAgents Agent Memory Open Source LLMs OpenAI · Anthropic · Google · Groq · Mistral · AWS · Microsoft · NVIDIA · Kimi · MiniMax · DeepSeek · Zhipu Alpha Vantage FRED · Polymarket StockTwits · Reddit · GPT-5.6 · GPT 5.5 · GPT-5.4 · Claude Sonnet 5 · Claude 4.6 · Fable 5 · Gemini 3.1 Grok 4.x · Qwen · GLM
  • 2026-09-10: Kernel-Managed Shared Memory for Multi-Agent System Personalization (breakingnewsofficial) · arxiv.orgAgent Memory Agents Context Engineering OpenAI · Meta · Mem0 · GPT-4o · Llama-3.1:8B · Qwen 2.5 7B
  • 2026-09-10: Research Attention Prediction: Evaluating LLM Agents’ Ability to Forecast AI Research Trends (breakingnewsofficial) · arxiv.orgAgents Agent Memory LLM Evals OpenAI · GPT 5.5 · Qwen3-4B
  • 2026-09-10: Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records (breakingnewsofficial) · arxiv.orgEmbeddings OpenAI · Anthropic · Google Crossref · Claude Sonnet 4.6 · GPT-5.2 · Gemini 2.5 Flash Sentence Transformers · Gemma · Qwen3
  • 2026-09-10: OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows (breakingnewsofficial) · arxiv.orgAgents Tool Use LLM Evals OpenAI · Anthropic · Google · Mistral AI · Microsoft · GPT-5.4 · Claude Opus 4.6 · Gemini-3.1-Pro · Qwen2.5-7B · Mistral-7B-v0.3 Phi-3.5-mini · Qwen2.5-1.5B
  • 2026-09-11: Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu (breakingnewsofficial) · arxiv.orgLLM Evals OpenAI · DeepSeek · GPT-5.1 Qwen-3-Max DeepSeek-3.1

FAQ

What is Alibaba?

Alibaba is a Chinese multinational technology conglomerate whose cloud division develops the open-weight Qwen model family and offers AI infrastructure through Alibaba Cloud. Its open releases are influential across the global open-model ecosystem. GROUNDING tracks Alibaba’s Qwen models, cloud AI services, and open-weight strategy.

What does this page track?

Dated radar mentions, source links, related concepts, and builder-relevant context for Alibaba, collected automatically by GROUNDING.

When was Alibaba last mentioned?

Alibaba was most recently mentioned in a radar update dated 2026-09-11.