The Real Story Behind Qwen3.8-Max’s AI Performance And Its Place In The Race

📊 Full opportunity report: The Real Story Behind Qwen3.8-Max’s AI Performance And Its Place In The Race on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has publicly confirmed its largest AI model, Qwen3.8-Max, with 2.4 trillion parameters and strong benchmark results. The open weights will be released next week, signaling a significant step in open AI development.

Alibaba has officially announced its flagship AI model, Qwen3.8-Max, with 2.4 trillion parameters and comprehensive benchmark results, after weeks of speculation and stealth preview. This marks a significant milestone in the company’s AI development and its push for open-weight models, which could influence the broader AI landscape.

On 3 August, Alibaba revealed the full specifications of Qwen3.8-Max, confirming it as a 2.4 trillion-parameter model built on the Qwen3.5 architecture, utilizing sparse mixture-of-experts technology. The model features a 95-billion active parameter count per query and supports multimodal input, including text, images, and video, with text output.

The benchmark table, published alongside the announcement, shows Qwen3.8-Max outperforming several competitors in key tests: achieving 86.6 on Terminal-Bench 2.1, surpassing Claude Fable 5 and Claude Opus 4.8, and only trailing GPT-5.6 Sol at 88.8. It also tops PaperBench at 93.0 and excels in multimodal and agentic tasks, notably achieving 86.1 on OSWorld-Verified and 91.5 on Parametric CAD Bench.

Alibaba confirmed the open weights will ship next week, with the 2.4 trillion-parameter checkpoint being a multi-node data center artifact. The company also announced a smaller 27 billion-parameter version, Qwen3.8-27B, optimized for local deployment on high-memory machines, which will be available sooner for practical use cases.

At a glance
updateWhen: announced August 3, 2023; details revea…
The developmentAlibaba announced the official details and benchmark results of Qwen3.8-Max, confirming its 2.4 trillion parameters and upcoming open-weight release.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Open-Weight AI Milestone

This announcement signifies a major step in AI transparency and accessibility, as Alibaba plans to release the largest open-weight model to date. The detailed benchmark results demonstrate competitive performance, especially in multimodal and agentic tasks, which could accelerate innovation and deployment in AI applications. The availability of the open weights next week will enable researchers and developers to experiment with a cutting-edge model, potentially shifting the landscape of open AI models.

However, the model's large size and complexity mean it remains primarily a data center artifact, limiting immediate widespread use. The smaller 27B version offers a more practical option for local deployment, but its performance relative to the flagship remains to be seen, especially regarding agentic capabilities after compression.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Development Timeline of Qwen3.8-Max

Alibaba's AI model development has been characterized by stealth and strategic releases. In July, the company previewed Qwen3.8-Max as an anonymous model called 'kaleb,' which was later confirmed during the World AI Conference in Shanghai. Prior to this, Alibaba's models, including Kimi K3 with 2.8 trillion parameters, had generated significant industry attention, but details were often withheld until official announcements.

The company's approach involved staged disclosures, culminating in the full benchmark release on 3 August. The model's architecture is based on Qwen3.5, utilizing sparse mixture-of-experts technology, and has demonstrated strong performance in various benchmarks, especially in multimodal and agentic tasks. The release of open weights aligns with broader industry trends towards transparency and democratization of AI tools.

Previous models like Kimi K3 and the anonymous 'kaleb' served as testing grounds for Alibaba's scaling and agentic capabilities, with the latest iteration showing significant improvements in long-horizon reasoning and environment interaction.

"Next week, we will release the open weights of Qwen3.8-Max, making the largest open-weight model publicly accessible for research and development."

— Alibaba spokesperson

Amazon

high-memory GPU for AI development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Deployment and Capabilities

It is still unclear how the open weights will be licensed and what restrictions, if any, will apply. The performance of the 27B version in real-world deployment remains untested, especially regarding agentic capabilities after compression. Additionally, the full implications of the model's size and architecture on accessibility and practical use are yet to be seen, as the 2.4 trillion-parameter checkpoint is inherently a multi-node artifact.

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

  • High-Performance Dual-Core CPU: Equipped with dual-core processor and Type-C USB
  • Rich Peripheral Support: Includes SPI, LCD, Camera, UART, I2C, and more
  • Wireless Connectivity: Supports Wi-Fi 2.4 GHz and Bluetooth 5 (LE)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba's Open-Weight AI Strategy

Alibaba plans to release the open weights of Qwen3.8-Max next week, enabling researchers and developers to evaluate its performance firsthand. The company is expected to publish licensing details and usage guidelines shortly thereafter. Meanwhile, the AI community will likely scrutinize the smaller Qwen3.8-27B model for practical deployment, and further benchmark results may emerge as the model is tested in diverse environments.

Industry observers will watch for how Alibaba's open model influences market competition and whether other firms follow suit in transparency and open deployment strategies.

Amazon

large-scale AI model deployment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main specifications of Qwen3.8-Max?

Qwen3.8-Max has 2.4 trillion parameters, built on the Qwen3.5 architecture, utilizing sparse mixture-of-experts, supporting multimodal input, with a 95B active parameter count per query.

When will the open weights be available?

Alibaba announced that the open weights of Qwen3.8-Max will ship next week, with the checkpoint being a multi-node data center artifact.

How does Qwen3.8-Max compare to competitors?

In benchmark tests, Qwen3.8-Max outperformed Claude Fable 5 and Claude Opus 4.8 in several tasks, and was only slightly behind GPT-5.6 Sol at maximum effort, showing competitive performance especially in multimodal and agentic benchmarks.

What is the significance of the smaller Qwen3.8-27B model?

The 27B version is designed for local deployment on high-memory machines and is expected to provide practical, accessible AI capabilities, though its performance relative to the flagship remains to be seen.

What are the potential limitations of this release?

The large size of the 2.4T checkpoint makes it a multi-node artifact, limiting immediate accessibility. Licensing details are still unpublished, and the impact of compression on agentic capabilities is uncertain.

Source: ThorstenMeyerAI.com

You May Also Like

The City That Watches Itself: The Living Digital Twin, and the God’s-Eye View We’re Building

Cities are developing dynamic digital twins integrated with AI and sensor technology, transforming urban planning and surveillance. But concerns over privacy and sovereignty persist.

The Anthropic-Blackstone-Goldman JV: Reverse-Engineering the $1.5B Enterprise AI Services Structure

Anthropic, Blackstone, and Goldman Sachs launch a new $1.5 billion standalone AI services company targeting mid-sized firms, embedding Anthropic engineers.

Spatial Focus Room: Make Distraction Impossible

A new deep-work app for Apple Vision Pro, Spatial Focus Room, removes distractions by creating immersive environments, transforming focus practices.

The Kill Switch: What the Anthropic Export Ban Really Costs the AI Industry

U.S. government’s export controls on Anthropic models led to immediate shutdowns, raising strategic and financial concerns for the AI industry amid ongoing disputes.