📊 Full opportunity report: Is Qwen3.8-Max The AI Second Place? The Numbers Tell A Complex Story on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has publicly released details of Qwen3.8-Max, claiming it as the second-largest AI model after Fable 5. Its benchmark results show strong performance, but only on selected tasks. The open-weight release next week will clarify its practical impact.

Alibaba has confirmed the release of Qwen3.8-Max, a 2.4 trillion-parameter AI model, with full benchmark results published today. The company claims it is second only to GPT-5.6 in performance, marking a significant milestone in large-scale AI development. This announcement, following weeks of speculation, provides the first detailed look at the model’s capabilities and specifications, and the open weights will be available next week.

Alibaba’s Qwen3.8-Max features approximately 95 billion active parameters per query within its 2.4 trillion total parameters, utilizing a sparse mixture-of-experts architecture based on Qwen3.5. The model is multimodal, supporting text, images, and video inputs, with text output. The benchmark table, published today, shows the model achieving top scores in several tasks, such as Terminal-Bench 2.1 (86.6), outperforming Claude models and only trailing GPT-5.6 Sol at 88.8. It also leads in paper-based benchmarks like PaperBench at 93.0 and demonstrates strong agentic capabilities, notably improving long-horizon task performance, such as in DeepSWE, which jumped from 21.6 to 56.6.

However, the model trails significantly in some software engineering benchmarks, like SWE-bench Pro (67.7 vs. Fable 5’s 80.0), indicating that its claim to being second in overall performance is based on a selective subset of tasks. The open weights, set to be released next week, are intended for high-end hardware, not for individual self-hosting, and the licensing details remain unpublished, adding a layer of uncertainty.

At a glance
reportWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba officially released detailed specifications and benchmark results for Qwen3.8-Max, confirming its status as a 2.4 trillion-parameter model and revealing performance metrics across various benchmarks.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Model Performance Claims

The announcement of Qwen3.8-Max and its benchmark results mark a notable development in large-scale AI models, especially given Alibaba’s claim of being second only to GPT-5.6. This influences perceptions of China’s AI competitiveness and the potential for open-weight models to challenge dominant players like OpenAI. The detailed benchmark data provides transparency, but the selectivity of the metrics and the upcoming open-weight release will determine its real-world applicability and influence on AI deployment strategies.

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

LLM Systems Engineering: Training and Building Large Language Models – Engineering AI Models Through Fine-Tuning, Continued Pretraining, and From-Scratch Development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recent Large-Scale Model Launches and Benchmark Trends

Over the past month, major AI companies have announced or previewed models exceeding 2 trillion parameters, including Alibaba’s Kimi K3 and others. The industry has seen a surge in multimodal and agentic AI capabilities, with benchmarks like Terminal-Bench and PaperBench becoming key indicators. Alibaba’s approach, initially stealthy, culminated in today’s detailed release, following a pattern of strategic announcements and selective disclosure. The open weights for Qwen3.8-27B, a smaller, more deployable version, are scheduled for next week, emphasizing Alibaba’s focus on practical, hardware-friendly models.

"Qwen3.8-Max demonstrates our advancements in multimodal AI and agentic capabilities, setting a new standard for large models."

— Alibaba spokesperson

Limitations of Current Benchmark Data and Open-Weight Details

While Alibaba has published comprehensive benchmark results, the model’s performance on untested tasks remains unknown. The upcoming open weights are designed for high-end hardware, not for general self-hosting, and the licensing terms are still unpublished. It is also unclear how well the agentic improvements will hold up in real-world, long-term applications, or whether the model’s performance across all benchmarks is representative of its overall capabilities.

Upcoming Open-Weight Release and Industry Benchmark Comparisons

Next week, Alibaba will release the open weights for Qwen3.8-27B, enabling wider testing and deployment. Industry observers will scrutinize its performance on real-world tasks and compare it against other leading models like GPT-5.6 and Claude. Further benchmark releases and independent evaluations are expected to clarify whether Alibaba’s claims translate into practical, scalable AI solutions or if the performance gaps in specific tasks limit its overall standing.

Key Questions

What does 'second only to Fable 5' mean in Alibaba’s claims?

It indicates that Alibaba considers Qwen3.8-Max to be the second-best performing large language model based on their benchmark results, primarily in certain tasks like Terminal-Bench and PaperBench, but this is selective and does not encompass all evaluation metrics.

Will the open weights be usable for individual researchers?

No. The open weights for Qwen3.8-27B are designed for high-memory, multi-node hardware, not for self-hosting on standard consumer machines. Licensing details are still unpublished, which may impact accessibility.

How does Qwen3.8-Max compare to other models like GPT-5.6?

In benchmark tests, Qwen3.8-Max scores highly on several tasks but trails GPT-5.6 at the top end. Its strengths lie in multimodal and agentic capabilities, but it does not outperform GPT-5.6 across all benchmarks.

What are the main limitations of the current data?

The benchmark results are selective, and performance on untested tasks remains unknown. The model’s practical deployment and licensing details are still evolving, which limits full assessment of its capabilities.

Source: ThorstenMeyerAI.com

You May Also Like

Capability or Control: The European Enterprise AI Playbook for the AI Act Era

A detailed analysis of how European companies are navigating the AI Act, focusing on capability versus control, licensing, and infrastructure choices.