The Most Recent AI Metrics Of Qwen3.8-Max: What They Tell Us

📊 Full opportunity report: The Most Recent AI Metrics Of Qwen3.8-Max: What They Tell Us on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has publicly disclosed comprehensive benchmark data for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and outperforms several competitors on key tests. Open weights will be available next week, with a smaller 27B version also announced. This development highlights Alibaba’s progress in large-scale AI, but some claims remain selective and unverified.

Alibaba has officially published detailed benchmark results for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and demonstrating strong performance across multiple tests. This marks a significant milestone in the company’s AI development, as it moves from stealth preview to full disclosure, with open weights arriving next week. The release provides transparency on the model’s capabilities and performance, which has implications for AI industry benchmarks and open-source efforts.

On August 3, Alibaba confirmed that its Qwen3.8-Max model features approximately 2.4 trillion total parameters, with about 95 billion active parameters per query, utilizing sparse mixture-of-experts architecture based on Qwen3.5. The model is multimodal, supporting text, image, and video inputs, with text output. The benchmark table, run on Alibaba’s own evaluation harness, shows the model achieving top scores in several key tests, including Terminal-Bench 2.1 at 86.6, surpassing Claude models but trailing GPT-5.6 Sol at 88.8. It also leads in PaperBench at 93.0 and performs strongly in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5.

Alibaba’s benchmarks reveal the model’s strengths in long-horizon reasoning and agentic execution, with significant improvements over its predecessor, DeepSWE, which jumped from 21.6 to 56.6. However, the model still trails in software-engineering benchmarks like SWE-bench Pro and FrontierSWE, with gaps of 12-15 points compared to Fable 5. The company also announced a smaller, 27B parameter version, Qwen3.8-27B, optimized for deployment on single high-memory machines, with details on its performance still pending. The open weights for the 2.4T model are set to be released next week, signaling a move toward more accessible large-scale models.

At a glance
updateWhen: announced August 3, 2023; benchmarks re…
The developmentAlibaba announced the full benchmark results for its Qwen3.8-Max model, confirming its size, performance, and upcoming open weights, marking a significant step in AI model transparency.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Transparency

The release of detailed benchmark data confirms Alibaba's progress in developing a large, multimodal AI model with competitive performance. The disclosure of 2.4 trillion parameters and the inclusion of open weights represent a significant step toward transparency and open AI development. These benchmarks allow industry comparison and could influence future AI model standards. However, the selective nature of the benchmarks and the absence of full licensing details mean the full impact remains to be seen. The move also signals Alibaba's intent to position itself as a major player in large-scale AI, potentially affecting industry dynamics and open-source contributions.

Autel MaxiSYS Ultra S2 AI Scanner, Intelligent Topology 3, Multi-Point DVI

Autel MaxiSYS Ultra S2 AI Scanner, Intelligent Topology 3, Multi-Point DVI

  • AI Diagnosis Support: AI Assistant and Data-Driven Diagnostics
  • Advanced Topology Map: Smart 3.0 ECU Network Analysis
  • Multi-Point DVI System: Comprehensive Digital Vehicle Inspection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's AI Model Development

Alibaba has been building its AI capabilities quietly since previewing Qwen3.8-Max in July, initially revealing only limited details and claiming it was 'second only to Fable 5.' The model was briefly introduced as 'kaleb' during the World AI Conference in Shanghai, with the full benchmark data withheld until now. Prior to this, Alibaba's AI models, such as Kimi K3 and earlier versions of Qwen, gained attention for their scale but lacked transparency in performance metrics. The recent benchmark release follows a trend of major tech firms publishing detailed model evaluations, signaling a shift toward greater openness in the industry.

The model's architecture, based on sparse mixture-of-experts, and its multimodal capabilities reflect Alibaba's strategic focus on versatile AI systems. The upcoming open weights and the smaller 27B version aim to facilitate wider adoption and deployment, especially for local-first applications. The benchmarks published today provide a clearer picture of where Alibaba's AI stands relative to competitors like OpenAI and other Chinese AI labs.

"We are committed to transparency and open collaboration, and the upcoming release of open weights for Qwen3.8-Max will enable broader innovation."

— Alibaba spokesperson

Etekcity Digital Body Weight Bathroom Scale, 440 lb Extra Wide Platform

Etekcity Digital Body Weight Bathroom Scale, 440 lb Extra Wide Platform

  • Large Platform Size: 13.8 x 11.8 inches with LCD display
  • High Weight Capacity: Supports up to 440 pounds
  • Accurate Measurements: High-precision sensors for reliable readings

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Licensing and Performance

Details about the licensing terms for the 2.4T open weights remain unpublished, raising questions about usage rights and commercial deployment. It is also unclear whether the benchmark performance on certain tasks, such as software engineering, fully reflects the model's capabilities or if further tuning is planned. Additionally, how the smaller 27B version will perform in real-world applications is still to be seen, as benchmark data for it has not yet been released. The long-term impact of these developments depends on licensing clarity and actual deployment performance.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice Capabilities: Camera and audio for AI interactions
  • Supports OpenCV & YOLO: Face tracking and human pose estimation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps: Open Weights and Broader Industry Impact

The immediate next step is the release of the 2.4 trillion parameter weights next week, which will allow third-party researchers and developers to evaluate and deploy the model independently. Alibaba is also expected to publish detailed licensing information, clarifying usage rights. The release of the 27B checkpoint will likely follow, targeting local deployment scenarios. Industry analysts will monitor how Alibaba's benchmarks influence competitors and whether the model's agentic and multimodal strengths translate into real-world applications. Further benchmarking and licensing details will shape the model's adoption and impact.

Building Robust AI Evals: Proven Strategies for Testing, Monitoring, and Improving LLM Performance (Engineered: Data, AI, and DevOps)

Building Robust AI Evals: Proven Strategies for Testing, Monitoring, and Improving LLM Performance (Engineered: Data, AI, and DevOps)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key capabilities of Alibaba's Qwen3.8-Max?

Qwen3.8-Max is a multimodal model supporting text, image, and video inputs, with high performance in reasoning, agentic tasks, and multimodal benchmarks, thanks to its 2.4 trillion parameters.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled to be released next week, enabling broader access for researchers and developers.

How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?

In benchmark tests, Qwen3.8-Max scores higher than Claude models but trails GPT-5.6 at maximum effort. It outperforms Fable 5 in some agentic and multimodal tasks but lags in software-engineering benchmarks.

What are the licensing implications of Alibaba's open weights?

The licensing details are not yet published, but historically Alibaba's open models have used Apache 2.0, while the new 2.4T checkpoint may have different terms, affecting usage rights.

What does this development mean for the AI industry?

This marks a move toward greater transparency and openness among major AI developers, potentially setting new standards for benchmark disclosure and model sharing, with industry-wide implications.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The Switch: You Never Owned the AI You Depend On

A government and companies can shut down AI models at any time, exposing dependencies on access rather than ownership. This impacts AI reliance and security.

Should You Trust Mistral Forge For Your AI Needs?

An analysis of Mistral Forge’s capabilities, ideal use cases, and when it may or may not be suitable for enterprise AI projects.

The Roblox Cheat That Broke Vercel.

A Roblox auto-farm script downloaded by an employee led to a major security breach at Vercel, exposing customer credentials across multiple cloud platforms.

AI Is the Alibi. The Reorg Is the Signal.

Coinbase’s recent layoffs and restructuring are officially linked to AI, but evidence suggests market pressures and crypto downturns are the primary drivers. Here’s what is confirmed and what remains uncertain.