Content Benefit

Content Strategy | UX Writing | Content Design | Globalization | Communications

Alibaba vs. Claude and Chat GPT

Alibaba vs. Claude and Chat GPT

In August came news that had been circulating in specialist forums for weeks: Alibaba unveiled Qwen3.8-Max, the most powerful model the Chinese giant has ever built, with the declared ambition of competing head-to-head with Claude and GPT. I use Claude every day, for work and to write this very blog, so the question hits close to home: is it worth switching?

What changes with Qwen3.8-Max, Alibaba’s new model

Qwen3.8-Max packs 2.4 trillion parameters, a figure that puts it close to Moonshot AI’s Kimi K3 (2.8 trillion) and well beyond most Western competitors, who usually keep this number under wraps. According to Alibaba, the AI can code for hours without supervision: in one internal trial it reportedly spent sixteen days developing and refining a coding tool on its own, an autonomy that echoes the themes covered in SaaS, AI and Vibe coding. On the benchmarks the company released, the scores are striking: 92.6 on GPQA Diamond, 86.6 on Terminal-Bench 2.1, 93 on PaperBench, and 86.1 on OSWorld-Verified. Record-breaking figures, with one important caveat: they all come from Alibaba itself, with no independent verification.

The real battleground: price

Here’s the part that actually matters to anyone weighing a move to another provider. Qwen3.8-Max costs around $2 per million input tokens and $6 per million output tokens, with discounted rates for previously processed content. Claude’s flagship version, by contrast, runs about $10 for input and $50 for output: a gap that reaches an eighth of the cost on the heavier side, generation itself. On paper, a difference like that isn’t trivial for anyone using artificial intelligence intensively, say for large volumes of content or code.

The flip side: when the model makes things up

There’s a catch that tempers the enthusiasm. According to the independent AA-Omniscience evaluation, cited by Tom’s Hardware, Qwen3.8-Max’s rate of fabricated answers jumped from 23 to 40 percent compared with the previous version: almost double. In practice, the assistant would rather guess than admit it doesn’t know, a habit that matters a great deal when a task gets delegated without checking every single step. On the Artificial Analysis Intelligence Index, the benchmark that ranks the various competitors, Qwen3.8-Max would score roughly in line with Claude Opus, but with one crucial difference: Anthropic’s results have been validated by outside reviewers, while Alibaba’s remain, for now, self-reported claims.

Why I’m not switching assistants, at least for now

I’ve spent months teaching Claude how I write: commas instead of long dashes, words to avoid, the ironic tone, never over the top, that I want for this blog. The list of rules I’m following right now, while writing this very piece, is the result of that work. Starting over with a different assistant would mean reloading every instruction, testing outputs from scratch, correcting things for weeks until the tone feels right again. It isn’t an insurmountable obstacle: it’s still an undertaking not to take lightly just because someone promises better benchmarks.

What would actually tempt me is one precise combination: results genuinely comparable, verified by third parties rather than just claimed, together with real and lasting savings. As long as one of those two pieces is missing, switching doesn’t pay off. And there’s a factor that often gets overlooked: artificial intelligence builds habits too, much like a CRM or an email client people stick with out of inertia more than conviction. I wrote about this when covering the struggles people run into using AI day to day: the learning curve and the time invested weigh as much as the model’s raw capabilities.

The bigger picture: a contest between China and the US

Qwen3.8-Max’s launch fits into the wider technology contest between Beijing and Washington, on top of being a commercial move: the same dynamic I already covered writing about Chinese marketplaces under fire. The United States keeps tight export controls on advanced chips, and Treasury Secretary Scott Bessent has promised zero tolerance in checking whether Chinese systems gained their capabilities through the unauthorized use of American technology. In this landscape, every recent release from Alibaba, Moonshot, or other Asian players should be read as a geopolitical statement as much as a product launch.

For now I’m sticking with Claude, but I’ll be watching the next round of independent evaluations on Qwen3.8-Max closely. If the promise of comparable results at such a reduced cost holds up under scrutiny from people with no stake in inflating the numbers, the question will become legitimate again. Until then, loyalty built over months of work is worth more than an announcement.

Related sources