Two weeks ago, Qwen3.8-Max existed only as two posts on X and a paid preview endpoint under the codename "Kaleb." No model card, no technical report, not a single published benchmark table. On August 3, Alibaba shipped the model for real: full API pricing, concrete rate limits, and its own benchmark table against GPT-5.6 Sol, Claude Fable 5, and Claude Opus 4.8. Some of those numbers look genuinely good for Qwen. One of them is built in a way that inflates the apparent progress over its predecessor.
What changed since the July 19 preview
When we covered the first Qwen3.8-Max announcement, two days after Moonshot AI's Kimi K3 launch, the list of unknowns was longer than the list of facts: no active parameter count, no pricing, not one published benchmark, and an open-weight release date that amounted to a rumor about "late July." August 3 closes most of those questions at once. Pricing is public, the model runs through a standard API, and Alibaba finally published a comparison table instead of just claiming "second-best AI in the world."
It doesn't close the most important question from two weeks ago, though: what it actually costs to run this model yourself. Without a disclosed active-parameter count, nobody outside Alibaba can calculate the real inference cost, only the price the company chose to charge for it.
The Moonshot AI thread hasn't gone away either. Alibaba still owns roughly 36% of the company whose model, Kimi K3, triggered this whole sequence of announcements. The full Qwen3.8-Max launch isn't a real-time reply anymore, the way the July 19 preview was, filed two days after Kimi K3. This is a separate stage: a company that holds a stake in its rival shipping its own product with full pricing and real numbers, regardless of what's happening on the other side of that same ownership structure.
Pricing that actually says something
$2 per million input tokens and $6 per million output is one flat rate across the entire 1-million-token context window: sending 5,000 tokens costs the same per token as sending 900,000. Cheaper cache tiers exist on top of that: 25 cents per million tokens read from cache, $2.50 to create an explicit cache, 17 cents to read from it. The rate limits are high for a model in this class, 2 million tokens per minute and 15,000 requests per minute, which suggests Alibaba actually wants people building production applications on this, not just running demos.
The model runs in xhigh reasoning mode by default, and tokens generated in that mode count as output tokens, billed at the higher rate. That's an easy detail to miss on a first API bill.
Benchmarks: strong where Alibaba picked the fight
On PaperBench and IFBench, Qwen3.8-Max genuinely leads, and by a real margin, not a fraction of a point. On Terminal-Bench 2.1 it loses to GPT-5.6 Sol but beats both Anthropic models. Outside the table, things get less comfortable for Alibaba: on SWE-bench Pro, Qwen scores 67.7 against Claude Fable 5's 80.0; on FrontierSWE, 73.5 against 88.8; and on Humanity's Last Exam, the model lands last among the four flagships compared, at 43.6 against Fable 5's 53.3. On GPQA Diamond, the 92.6 score is barely above the previous version's 92.4 for Qwen3.7-Max.
Every one of these numbers comes from material Alibaba itself published. No independent benchmarking organization, including Artificial Analysis, had published its own measurement of Qwen3.8-Max at launch time. For comparison, the predecessor, Qwen3.7-Max, saw a sharp drop between benchmark methodology versions in Artificial Analysis's own Intelligence Index: 56.6 in the May v4.0 methodology, 46 in the July v4.1 revision. A good reminder that the measurement methodology itself can move a score more than the next model version does, so it's worth waiting for independent numbers before anyone calls a final ranking.
The pattern in the results is fairly readable: wherever the task is about precisely following complex instructions (IFBench) or summarizing and analyzing long documents (PaperBench), Qwen3.8-Max genuinely leads by a clear margin. Wherever the task is about autonomous work on real code over a longer stretch, SWE-bench Pro and FrontierSWE, the lead flips to Claude Fable 5, and not by a few points, by more than ten. For anyone looking for a model to run unsupervised coding agents, that gap matters more than the overall ranking does.
Open weights: the promise slipped by a week
In July, the only information about open weights was a rumor about "late July," never confirmed by Alibaba. That date has already passed, and the company now says "next week" counting from the August 3 launch, which lands realistically around mid-August. Two variants are promised: the full 2.4 trillion parameters, requiring a multi-node cluster (the weights alone at 4-bit precision would take up roughly 1.2 terabytes of memory), and Qwen3.8-27B, built as a version that can run on a single GPU server. This is the second promised, and still unmet, date in this model's history, worth remembering when weighing Alibaba's next promises.





Comments
Discussion
Join the conversation around this story.
Join the discussion
Sign in to comment and reply to other readers.
Sign inNo comments yet
Start the discussion first.