Anyone working with AI this summer has two names on the table: GLM 5.3 from Z.ai, arrived on August 14 with open weights, and Claude Fable 5 from Anthropic, the reasoning model released in June that many independent benchmarks place at the top. They are the two text models of the moment, and the question I ask myself is the same as anyone producing something with these tools: is it worth paying ten times as much? My answer is no, and the numbers to say it are all public. Let us look at them together.
What GLM 5.3 is, in thirty seconds
GLM 5.3 is the flagship model of Z.ai, the lab born as a spin-off of Tsinghua University in Beijing, formerly Zhipu AI, now listed on the Hong Kong stock exchange. It came out on August 14, 2026 and has one technical quirk: under the hood it is the same 743-billion-parameter base model as GLM-5.2, improved only with post-training. The context window reaches one million tokens, reasoning is always on with three intensity levels, and the model works on text only: no images, no video.
Two things matter to whoever pays the bills. First: the weights are open, with a license that allows free commercial use for companies under 10 billion dollars in revenue, which means practically all of them. Second: the lighter Flash version runs on promo at 0.075 dollars in input and 0.25 in output per million tokens, under the MIT license.
The table that matters
Now the comparison. The table pits the two models against each other on seven benchmarks: software development on real projects (FrontierSWE), offensive security (ExploitBench), the hard Humanity’s Last Exam with tools, two versions of Terminal Bench, defensive cybersecurity (CyberGym) and Z.ai’s coding suite.
The result says 5 to 2 for Fable. But scores do not tell you the distance: on Humanity’s Last Exam the two models are separated by 1.4 points, on Terminal Bench 2.1 by 0.2, with GLM ahead. Where Fable really pulls away we will see in a moment. The question that matters to producers, though, is another: do those points justify a bill ten times higher?
Where Fable 5 stays ahead
Fable 5 wins where the hardest software development work is measured. On FrontierSWE, real projects with repositories and tests, it leads by ten points: 88.2 against 78.1. On Terminal Bench 3.0, the latest version of the terminal test, 33.7 against 28.3. On Z.ai’s coding suite, 39.5 against 31.4.
Then there is ExploitBench, the largest gap in the whole table: 78.0 against 54.4, almost twenty-four points. It is the benchmark measuring the ability to find and exploit security vulnerabilities: security research territory, not everyday use. If your job is spending hours inside huge repositories with agents that must fend for themselves, this gap exists and it is real: Fable 5 is the champion of the category.
Where GLM 5.3 wins and where it ties
GLM 5.3 wins the two tests where the security game is played. On CyberGym, the defensive cybersecurity benchmark, it scores 84.5 against 83.8: the highest score among all open models. And it is not a slide-deck number: in the public audit Z.ai logged 2,436 vulnerabilities found in real software, each verified and tracked in an open register anyone can check.
On Terminal Bench 2.1 it flips the result: 88.2 against 88.0. On Humanity’s Last Exam with tools it stays 1.4 points from the top. And in the same publication GLM 5.3 beats Claude Opus 4.8 on three of the seven tests: Terminal Bench 2.1, Terminal Bench 3.0 and CyberGym. The picture is that of a top-tier model, with its own peaks in security.
The price: seven times less going in, eleven coming out
Now the part that decides everything for me. Fable 5 costs 10 dollars per million input tokens and 50 in output: double Opus 4.8 at 5 and 25, same price list as Opus 5. GLM 5.3 costs 1.40 and 4.40. Seven times less on input, eleven times less on output.
A concrete example: a month of work with agents consuming 20 million tokens in and 10 million out costs 700 dollars with Fable 5 and 72 with GLM 5.3. That is not a margin, it is an order of magnitude.
Two honest notes. Anthropic applies a 90 per cent discount on input with prompt caching, so anyone repeating the same context many times sees the input gap shrink a lot; output, though, the line that grows with agents reasoning at length, stays eleven times more expensive. And since July Fable 5 is no longer included in Anthropic’s Max and Pro plans: it remains available only on the API, pay per credit.
The Ox Alpha story
If it looks to you like price does not matter much, let me tell you how the most talked-about bet of the summer ended. In early August on OpenRouter, the platform where you use models on consumption, a nameless model appeared, code name Ox Alpha, with very low prices.
In three days it consumed 11 trillion tokens, becoming one of the most used models on the platform. On August 26 Z.ai revealed it was theirs: behind the pseudonym was GLM 5.3. On the stock exchange the group’s share jumped 12 per cent in a day, and Artificial Analysis places the model on par with Kimi K3 among the best open ones around.
The moral is not that GLM 5.3 is unbeatable: it is that at equal tier, whoever sets the right price wins. That is exactly the calculation I make when choosing my studio’s tools.
Why I prefer GLM 5.3
My preference, in one sentence: I prefer GLM 5.3 over Fable 5 because it costs far less and the results are very similar. It is not fandom, it is a producer’s calculation. In my work, mini-films, commercials and visuals, the big budgets live in video models and generation platforms; the text and reasoning part is for scripts, production documentation, complex prompts and automations. Real workloads, but ones where a handful of benchmark points does not justify a bill ten times higher.
There is also a reason of independence. Open weights mean you can take the model where you need it and run it where it is convenient, without depending on anyone’s pricing decisions. The July lesson, when Fable 5 left the subscription plans overnight, applies to everyone: when the vendor changes the terms, whoever holds the weights still has a model.
If you want to know Fable 5 in detail, I told its story in this article. And the final advice is the same I give for video tools: try both on your real case, with your prompts and your workloads. Benchmarks tell you the tier, the price list tells you how long you can afford it. If instead your brand needs someone who uses these tools every day, that is literally my job.
Leave a comment
Comments are reviewed and approved before they appear.