Why the 1,100% DeepSeek headline does not match most bills
The friction for buyers is specific:
- Different outlets quoted different tiers. "11x", "over 1,100%", and "350%" are all accurate. They map to peak cache-hit input, output, and cache-miss input — not one blended rate.
- Peak hours are now priced in. Peak is 9am–12pm and 2pm–6pm Beijing time. The announcement's "more flexible workload scheduling" is capacity talk, not marketing fluff.
- Official is no longer always cheapest. At peak, DeepSeek's own API now sits above several resellers (GMI Cloud, Novita, and others still list V4 Pro below the new official peak).
- Open weights are not Apache by default. Qwen3.8-Max is downloadable, but the custom license keeps commercial leverage over large MaaS and assistant businesses.
We already covered DeepSeek V4 Flash GA scores and the old flat rate and the Qwen3.8-Max launch window. This piece is the mid-August three-lab shift, not another single-model recap.
What happened in two weeks: the timeline
From mid-July through 00:00 Beijing time on 17 August:
| Date | Event |
|---|---|
| 16 Jul 2026 | Moonshot AI open-weights Kimi K3 (2.8T parameters), drawing US security scrutiny |
| 2–3 Aug 2026 | Alibaba previews, then launches, Qwen3.8-Max as a hosted API |
| 10 Aug 2026 | Meta releases Muse Glimmer (30B, Apache 2.0), teases open weights for flagship Muse Spark 1.2 |
| 12 Aug 2026 | Alibaba publishes Qwen3.8-2.4T-A95B on Hugging Face / ModelScope; xAI ships Grok 4.6 |
| 13 Aug 2026 | DeepSeek-V4-Pro goes GA and announces a price increase effective 17 Aug; Google ships discounted Gemini 3.7 Flash |
| 14 Aug 2026 | Zhipu ships GLM-5.3, reusing GLM-5.2's 743B base |
| 17 Aug 2026, 00:00 Beijing | DeepSeek's new pricing takes effect |
Zoom out: on 30 Jul, OpenAI cut GPT-5.6 Luna 80%, then on 6–7 Aug made Luna the free default with unlimited text chats. Chinese labs raised prices and opened flagship weights while US labs cut prices and went free at the consumer layer. Same fight, two sides. The Luna cut is unpacked in our GPT-5.6 pricing note.
Rate cards and specs: is DeepSeek still the cheapest frontier-class API?
DeepSeek rates below are RMB per 1M tokens, effective 17 Aug 00:00 Beijing time. Peak hours: 9am–12pm and 2pm–6pm Beijing.
| Billing item | Old | New off-peak | New peak | Peak increase |
|---|---|---|---|---|
| V4-Flash cache hit (input) | ¥0.02 | ¥0.05 | ¥0.10 | ~400% |
| V4-Flash cache miss (input) | ¥1.0 | ¥1.5 | ¥3.0 | 200% |
| V4-Flash output | ¥2.0 | ¥4.5 | ¥9.0 | 350% |
| V4-Pro cache hit (input) | ¥0.025 | ¥0.15 | ¥0.30 | ~1,100% |
| V4-Pro cache miss (input) | ¥3.0 | ¥4.5 | ¥9.0 | 200% |
| V4-Pro output | ¥6.0 | ¥13.5 | ¥27.0 | 350% |
The 1,100% figure applies to peak-hour cache-hit input — the tier that started closest to free. Output, which dominates most real bills, rose 350%. Independent cost modeling found a realistic heavy-usage workload (roughly 84M tokens/month, mostly off-peak, half cache hits) sees a bill increase closer to 1.8x.
Qwen3.8-2.4T-A95B (Qwen3.8-Max open weights):
| Spec | Detail |
|---|---|
| Parameters | 2.4T total, 95B active per token (MoE, 512 experts, 10 routed + 1 shared) |
| Context window | 262,144 tokens native (open checkpoint), extendable to ~1.01M; hosted Max defaults to 1M |
| Release cadence | Preview 2 Aug → API live 3 Aug → open weights 12 Aug |
| API pricing (international) | $2/M input, $6/M output |
| License | Not Apache 2.0 — a custom Qwen3.8-Max License |
| Why it matters | First time Alibaba has open-weighted a Max-tier model; Qwen3.5/3.6/3.7 Max stayed API-only |
GLM-5.3 vs GLM-5.2: same base, post-training only. These are Zhipu's own reported numbers — no independent third-party re-run has been published yet.
| Benchmark | GLM-5.2 | GLM-5.3 | Change |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6% | 28.3% | +23.7 pts |
| DeepSWE v1.1 | 46.2% | 66.9% | +20.7 pts |
| Agents' Last Exam (CLI) | 23.8% | 28.5% | +4.7 pts |
| CyberGym | 77.2% | 84.5% | +7.3 pts |
| AutomationBench | 26.2% | 48.2% | +22.0 pts |
GLM-5.3 still trails GPT-5.6 Sol (34.6%) and Claude Fable 5 (33.7%) on Terminal-Bench 3.0. It is a top open-weight result, not an outright frontier win.
Head-to-head (per 1M tokens; RMB-to-USD at ~¥7.15/$1, approximate):
| Model | Input | Output | Open weights? |
|---|---|---|---|
| DeepSeek V4-Pro (peak) | ¥9.0 (~$1.26) | ¥27.0 (~$3.78) | No |
| DeepSeek V4-Pro (off-peak) | ¥4.5 (~$0.63) | ¥13.5 (~$1.89) | No |
| Qwen3.8-Max (international API) | $2.00 | $6.00 | Yes (custom license) |
| OpenAI GPT-5.6 Luna | $0.20 | $1.20 | No |
| Claude Opus 5 (implied, per Alibaba's comparison ratio) | ~$5.00 | ~$25.00 | No |
Short answer: no. Off-peak V4-Pro is still well below Claude Opus 5, but it is no longer the outright cheapest option. Qwen3.8-Max international pricing and OpenAI Luna both undercut DeepSeek's off-peak rate. "Chinese model = cheapest model" held for most of 2025 and early 2026. It is not a safe assumption anymore.
Three strategies: capacity pricing, licensed open weights, post-training scale
DeepSeek: from flat-rate cheap to time-of-day pricing. The easy misread is "China's cheapest model caved to margin pressure." The structure reads more like compute constraints made visible. Flat, always-cheap pricing worked as acquisition while GPU supply kept up. Once usage grew exponentially and capacity did not, something had to become explicit. "Encouraging more flexible workload scheduling" is corporate-speak for scarce peak-hour compute.
Alibaba: open weights buy mindshare; a custom license protects the revenue ceiling. The full 2.4T checkpoint is free to download. The license is not the permissive Apache 2.0 used for smaller Qwen models. Any Model-as-a-Service or AI Work Assistant business earning over $50 million in any 12-month period must negotiate a separate commercial license. Products with 100M+ monthly active users or $20M+ in monthly revenue must prominently display the model name. That is a different bet than Meta's Muse Glimmer under unrestricted Apache 2.0.
One rumor to kill: claims that Alibaba's license bans downloads from the US, EU, UK, and South Korea. False. The published license text contains no geographic or territorial clause.
GLM-5.3: no new base model, a bigger post-training bet. Same 743B-parameter base as GLM-5.2, no retraining, and a roughly 6x jump on Terminal-Bench 3.0 (4.6% → 28.3%) from scaling reinforcement-learning environments. As pretraining scaling laws show diminishing returns, post-training RL scale is becoming an independent lever with a lower cost floor than retraining a foundation model. Mid-tier labs can still close the gap on agentic and coding benches.
Six steps to reconcile the headline with your invoice and license
Pricing, license terms, and scores change. Verify the official pages before you republish or reprice a product.
- Copy the official rate card, not a screenshot of a headline. DeepSeek splits cache hit, cache miss, and output. Write down the rows your traffic actually hits.
- Map your local clock to Beijing peak windows. Peak is 09:00–12:00 and 14:00–18:00 Asia/Shanghai. Overnight in one region can be midday in another.
- Estimate cache-hit rate from last week's logs. Long fixed system prompts absorb the cache-hit hike. Low hit rate means the 350% output and 200% miss-input rows dominate.
- Read the Qwen license file, not the announcement thread. Personal and internal use are largely unaffected. MaaS or AI Work Assistant revenue over $50M in any 12 months needs a separate commercial license.
- Book GLM-5.3 scores as vendor-reported until a third party re-runs them. Terminal-Bench 3.0 still trails GPT-5.6 Sol and Claude Fable 5.
- If you download weights for a side-by-side, use a wipeable box. A 2.4T checkpoint is not a laptop job. For quantized smaller models or agent workflows on Apple Silicon, a dedicated, rootable machine beats a shared notebook.
TZ=Asia/Shanghai date '+%H:%M %Z'
python3 -c "from datetime import datetime; from zoneinfo import ZoneInfo
h=datetime.now(ZoneInfo('Asia/Shanghai')).hour
print('peak' if (9<=h<12 or 14<=h<18) else 'off-peak')"
Reseller list prices can undercut official peak rates for a while. Put supplier identity, quota, and cutoff risk on the same line as the sticker price.
What is disputed, and where to verify
- The 1,100% headline is technically accurate and still misleading without the tier. It is peak cache-hit input. Output — the cost that dominates most real bills — rose 350%.
- Claims that Qwen3.8-Max runs on Alibaba's in-house Zhenwu M890 chips (and "Pangu AL128" supernodes), reported by several Chinese financial outlets, have not been independently confirmed by Alibaba technical docs or third-party benchmarks. Treat as vendor-adjacent, unverified.
- GLM-5.3's reported discovery of a "serious vulnerability" in Cursor comes from VentureBeat and Zhipu's own disclosure. Technical details have not been made public. Vendor-sourced, not independently audited.
- Reports that China's Ministry of Commerce may prepare retaliatory export controls on AI/semiconductor technology are speculative, not an official announcement. Background context only.
Chinese financial media has started calling the last month "three model updates a week." US labs ran the opposite play at the consumer layer. There is also a geopolitical reading: Kimi K3 already drew US security scrutiny; Alibaba choosing this window for a 2.4T flagship has been read as locking in international mindshare before any regulatory tightening. That is an informed interpretation, not a confirmed fact.
Official sources (check the live page; figures may change):
DeepSeek API Docs — Models & Pricing
Hugging Face — Qwen/Qwen3.8-2.4T-A95B
Z.ai — GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
Third-party reporting:
CIO — DeepSeek raises some V4 prices by more than 10x
NVIDIA Technical Blog — Serve Qwen3.8-2.4T-A95B
FAQ
Is DeepSeek still cheaper than GPT-5.6 or Claude after the price hike?
Its off-peak rate is still cheaper than Claude Opus 5, but it is no longer the single cheapest option overall. OpenAI's GPT-5.6 Luna ($0.20/$1.20 per million tokens) and Alibaba's international Qwen3.8-Max pricing ($2/$6) now undercut DeepSeek's new off-peak rates on at least one dimension. DeepSeek is still relatively cheap for a frontier-class model, just not the outright cheapest anymore.
Can I use Alibaba's Qwen3.8-Max open weights for free in a commercial product?
Yes, for most use cases — personal projects and internal enterprise use are unaffected. The catch applies only if you are running a Model-as-a-Service or AI Work Assistant business that has earned over $50 million in any consecutive 12-month period; that tier requires a separate commercial license from Alibaba.
Is Qwen3.8-Max banned or restricted for US, EU, or UK users?
No. That claim circulated online but is false — the published license contains no geographic restriction of any kind. The restrictions are revenue-based (tied to how much money your service makes), not tied to where you or your users are located.
What's actually different between GLM-5.3 and GLM-5.2?
Nothing at the base-model level — both use the same 743-billion-parameter foundation model. The performance gains (roughly 6x on Terminal-Bench 3.0) come entirely from scaling up reinforcement learning during post-training, with no retraining of the base model.
Will Meta actually open-source its flagship model, not just the smaller Muse Glimmer?
Not yet. Muse Glimmer is a 30B distilled model, not Meta's real flagship. CEO Mark Zuckerberg has said open weights for the larger, closed Muse Spark 1.2 are coming "soon," which — if it happens — would make it the first US flagship-tier model released openly. As of this writing, that release has not happened; treat it as a stated intention, not a confirmed fact.
Time-of-day API pricing, a 2.4T download, and a post-training bake-off all hit the same constraint: a shared laptop is dirty, a local disk fills up, and you should not mix eval traffic with day-to-day compiles. When you need a clean, rootable Apple Silicon box you can wipe after the run, a KVMFLUX cloud Mac mini rented by the day is usually the cleaner path — dedicated hardware, SSH + VNC. Period math is in daily vs weekly vs monthly vs quarterly.
Compare open weights on a dedicated Mac, by the day
Physical Mac mini M4, root and SSH, for quantized models or agent workflows — without fighting a shared laptop for disk and state.