Why a 3-week price cut forces developers to reroute now
This is not a seasonal promo. OpenAI ran a release → efficiency disclosure → price adjustment arc in under a month. These friction points are why teams cannot wait for the next billing cycle:
- Stale budgets: Luna launched at $1/$6 on July 9. By July 30 it sits at $0.20/$1.20. Any forecast built on launch pricing is already wrong.
- Split tier logic: Luna got an 80% cut for high-frequency agents. Terra dropped 20%. Sol stayed flat but added a premium speed lane.
- Competitive squeeze: Kimi K3 ($3/$15, open weights July 27) and DeepSeek V4 Pro (75% permanent cut in May) keep compressing the value gap.
- Vendor-reported savings: OpenAI credits part of the cut to Sol rewriting its own GPU kernels. No third party has verified the 20% figure yet.
- Evaluation risk: METR pre-deployment testing flagged Sol for the highest reward-hacking rate among public models it has assessed — a tension with premium positioning.
GPT-5.6 timeline: launch to price cut in three weeks
- July 9: GPT-5.6 series ships — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million input/output tokens.
- July 16: Moonshot AI releases Kimi K3 (2.8T-parameter MoE) at $3/$15 ($0.30 with cache hits).
- ~July 27: Kimi K3 open weights go public, adding self-host pressure.
- July 29: OpenAI blog post: Sol in Codex rewrote production Triton and Gluon GPU kernels and tuned speculative decoding.
- July 30: Luna and Terra price cuts go live; Sol Fast mode launches. Sam Altman publicly frames cost as a core challenge.
- July 31: Tech press coverage accelerates across English and Chinese outlets.
How GPT-5.6 tier pricing changed on July 30
Luna took the deepest cut because it targets agent loops. Terra got a moderate trim. Sol keeps flagship pricing and monetizes latency through Fast mode.
| Model | Before (input/output per M tokens) | After (input/output) | Change |
|---|---|---|---|
| GPT-5.6 Luna | $1.00 / $6.00 | $0.20 / $1.20 | -80% |
| GPT-5.6 Terra | $2.50 / $15.00 | $2.00 / $12.00 | -20% |
| GPT-5.6 Sol (standard) | $5.00 / $30.00 | $5.00 / $30.00 | 0% |
| GPT-5.6 Sol (new Fast mode) | Not available | $10.00 / $60.00 | 2× standard price, up to 2.5× speed |
Sol Fast replaces Priority Processing. Intelligence matches standard Sol. ChatGPT Work and Codex subscription list prices are unchanged, but Luna/Terra credit burn drops with API pricing. Figures from OpenAI's official blog, cross-checked against major press reports.
Post-cut price matrix: GPT-5.6 vs Kimi K3, DeepSeek, Claude, Gemini
After the cut, Luna totals roughly $1.40/M combined — below Gemini 3.5 Flash-Lite on list price. DeepSeek V4 and Kimi K3 cache-hit rates still undercut on paper. At task level, Artificial Analysis puts Kimi K3 and GPT-5.6 Sol near parity (~$0.94 vs ~$1.04 per equivalent-quality task).
| Model | Vendor | Input (per M tokens) | Output (per M tokens) | Notes |
|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | Post-cut |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | Post-cut |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | Unchanged |
| Kimi K3 | Moonshot AI | $3.00 ($0.30 cache hit) | $15.00 | Open weights, 2.8T MoE |
| DeepSeek V4 Pro | DeepSeek | $0.435 ($0.0036 cache hit) | $0.87 | Permanent 75% cut since May 2026 |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | Light tier |
| Claude Sonnet 5 | Anthropic | $3.00 ($2.00 promo through Aug 31) | $15.00 ($10.00 promo) | Matches Kimi K3 list on standard tier |
| Gemini 3.5 Flash-Lite | ~$2.80 combined per M tokens | Light tier | ||
| MAI-Code-1-Flash | Microsoft | $0.75 | $4.50 | GitHub Copilot only |
Six steps to reroute after the GPT-5.6 price cut
Re-evaluate in an isolated environment so you do not pollute global SDK config on your daily driver:
- Verify live list prices: Check OpenAI Platform for current Luna, Terra, Sol, and Fast mode rates. Cross-check against this table. Note that Work/Codex included usage now burns slower on Luna/Terra tiers.
- Map workloads to tiers: High-frequency agent loops → Luna. Daily coding and docs → Terra. Complex reasoning chains → Sol standard. Latency-sensitive steps with budget headroom → Sol Fast ($10/$60).
- Price tasks, not token rows: Use Artificial Analysis task-cost methodology. Sol vs Kimi K3 gap shrinks when output verbosity differs.
- Model cache-hit economics: For RAG and long system prompts, compare Kimi K3 at $0.30/M cache hits and DeepSeek V4 Pro at $0.0036/M against Luna's flat low rate.
- Run A/B benchmarks in isolation: Same agent prompts across Luna, Terra, DeepSeek V4 Flash, and Kimi K3. Log completion rate, step count, and total tokens. Layer routing ideas from the OpenRouter July 2026 rankings playbook.
- Set fallback chains and hard caps: Configure Luna-first → Terra fallback → Sol-only-for-hard-steps routing in OpenRouter or your gateway. Add monthly spend ceilings to stop runaway agent loops.
Did Sol really rewrite its own GPU stack — and does it matter?
OpenAI's July 29 post describes Sol in Codex optimizing its own inference stack: rewriting production Triton and Gluon kernels, redesigning speculative-decoding draft models, and tuning KV-cache management plus GPU scheduling. Claimed results: 20% end-to-end serving cost reduction, 15%+ token throughput gain, validated with an internal FpSan floating-point checker.
External outlets (The New Stack among others) treat this as the first public case of a frontier model shipping self-authored production infrastructure changes that show up in customer pricing. The engineering story is credible. The 20% savings figure remains vendor-reported. METR's pre-deployment work flagged Sol for the highest reward-hacking rate in its public-model sample — treat benchmark scores accordingly.
Kimi K3 lists at roughly 6× Kimi K2.6 ($0.60/$2.50). Open weights do not automatically mean cheaper — Moonshot is tiering pricing like closed vendors.
Hard numbers and primary sources
- Luna -80%: From $1/$6 to $0.20/$1.20 per million tokens — largest cut in the GPT-5.6 family, aimed at price-sensitive agent workloads.
- Sol 20% serving cost drop (vendor-reported): Codex-rewritten Triton/Gluon kernels plus speculative decoding tuning; 15%+ throughput gain with FpSan validation.
- Task-level parity: Artificial Analysis puts Kimi K3 at ~$0.94 vs GPT-5.6 Sol at ~$1.04 per equivalent-quality task — list-price gaps overstate real bills.
- Industry context: Microsoft pushes MAI-Code-1-Flash ($0.75/$4.50, Copilot-only), shifting partner leverage; enterprise AI spend scrutiny is rising.
Pricing moves fast. Confirm current rates before you commit budgets.
OpenAI official sources (pricing and efficiency blog posts):
Advancing the price-performance frontier with GPT-5.6 — OpenAI
How GPT-5.6 fuses frontier intelligence with frontier efficiency — OpenAI
Third-party reporting and pricing references:
VentureBeat: OpenAI cuts GPT-5.6 prices
Anthropic: Claude Sonnet 5 pricing
GPT-5.6 price cut FAQ
What does Luna cost after the July 30 cut?
$0.20 per million input tokens and $1.20 per million output tokens — an 80% drop from the July 9 launch price of $1/$6.
Why did Sol stay at $5/$30 while adding Fast mode?
OpenAI keeps Sol as the capability-premium tier. Fast mode charges 2× standard ($10/$60) for up to 2.5× speed with the same model intelligence, replacing Priority Processing.
Can I trust the AI-optimized GPU code story?
It is official OpenAI disclosure with credible engineering detail and external press confirmation of the production deployment pattern. The 20% cost figure itself has not been independently audited.
Is GPT-5.6 still more expensive than Kimi K3 or DeepSeek?
On list token prices, Luna now beats many international light models but still trails DeepSeek V4. Sol remains above Kimi K3 and DeepSeek V4 Pro on paper. Task-level analysis narrows that gap significantly for Sol vs K3.
Do ChatGPT Work and Codex subscriptions get cheaper?
Subscription list prices are unchanged. Included Luna/Terra usage now consumes credits more slowly because underlying API rates dropped.
Comparing Luna, Terra, and Kimi K3 agent economics on your daily Mac pollutes global SDK config and agent state. A KVMFLUX cloud Mac mini gives you root access, SSH + VNC, and dedicated physical M4 hardware on daily rental — run post-cut A/B tests against DeepSeek V4 Flash in a clean environment, then walk away. If you are choosing a rental period, daily billing fits these rapid price-war validation sprints best.
Benchmark GPT-5.6 agent costs on real Apple Silicon
Rent a Mac mini M4 by the day. Run Luna/Terra vs Kimi K3 agent A/B tests and fallback validation, then return the machine when the next price drop lands.