What's confirmed vs. what's just Musk's word
Only one thing is on the record here: a tweet. Everything else — the 1.5 trillion parameters, the significantly improved SFT and RL, the follow-up Grok 4.7 at 2.1 trillion — comes from that single X post, with no accompanying model card, benchmark suite, or pricing page the way xAI published for Grok 4.5. That distinction matters for anyone deciding whether to build around this release.
- Single-source timeline: xAI's blog and product pages have not confirmed the August 7 date or parameter counts. Musk's public statements have a track record of slipping — "Musk time" is a real planning variable.
- Zero published benchmarks: Unlike Grok 4.5's launch with 15 tracked scores, Grok 4.6 has no independent evaluation data yet.
- Compressed evaluation windows: Grok 4.5 shipped July 8. A 4.6/4.7 double-drop in August means any 4.5 baseline you built has a shelf life measured in weeks, not quarters.
- Competitive pressure is real: Kimi K3's full open-weight release topped Frontend Code Arena at 1,679 points — the first open-weight model to beat every closed model on that board. Musk himself commented "impressive" on related benchmark posts.
- Vendor risk on a separate axis: In July 2026 xAI sued a user for allegedly using Grok to generate CSAM. A January 2026 Common Sense Media report rated Grok among the worst AI chatbots for child-safety risks. None of this affects technical capability, but enterprises should factor it into vendor evaluation.
The Grok 4.5 to 4.7 timeline
- July 8, 2026 — xAI ships Grok 4.5, built for coding and agentic work, co-trained with Cursor on real developer session data. 500K-token context, $2/$6 per million input/output tokens, published model card with 15 benchmark scores.
- July 16–26, 2026 — Moonshot AI's Kimi K3 goes from hosted preview to fully open-weight release (2.8T parameters, 1M-token context), immediately topping Hugging Face's trending chart.
- July 28, 2026 — Musk posts the Grok 4.6/4.7 roadmap in reply to Rauch. Same day, 1,200+ employees across OpenAI, Anthropic, Google DeepMind, and Meta publish the "Pacing the Frontier" letter.
- ~August 7, 2026 (target) — Grok 4.6, 1.5T parameters, positioned as an SFT/RL upgrade rather than a raw scale-up.
- Late August–early September 2026 (estimated) — Grok 4.7, 2.1T parameters. Musk says it will be "better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency."
xAI, Tesla, and SpaceX timelines from Musk have historically slipped by days to weeks. Treat "around August 7" as a target, not a guarantee — verify against official xAI channels before committing architecture decisions.
Grok lineup specs at a glance
| Model | Date | Parameters | Focus | Status |
|---|---|---|---|---|
| Grok 4.3 Beta | Apr 17, 2026 | Undisclosed | Baseline | Shipped |
| Grok 4.5 | Jul 8, 2026 | Undisclosed | Coding/agentic, Cursor co-training | Shipped, benchmarked |
| Grok 4.6 | ~Aug 7, 2026 | 1.5T | SFT/RL upgrade | Announced via tweet, unshipped |
| Grok 4.7 | ~late Aug–early Sep 2026 | 2.1T | Broad upgrade over 4.6, better token efficiency | Announced via tweet, unshipped |
All Grok 4.6/4.7 figures are unverified vendor claims from a single social media post — treat them as directional, not confirmed specs.
Why xAI is emphasizing post-training, not just scale
Supervised fine-tuning (SFT) trains a model on curated example outputs to shape its behavior; reinforcement learning (RL) uses reward signals to teach a model which action sequences actually work — critical for multi-step agentic tasks. Musk's framing — "significantly improved SFT & RL" — signals xAI is doubling down on the same playbook that made Grok 4.5 competitive on agentic benchmarks: Grok 4.5 used roughly 15,954 output tokens per SWE-Bench Pro task versus Opus 4.8's 67,020, a 4.2x efficiency gap largely credited to post-training on real Cursor developer sessions.
Grok 4.6's jump to 1.5T parameters is a real scale increase, but Musk's own framing of Grok 4.7 — bigger at 2.1T, "better in every way except slightly slower to serve" — suggests xAI is deliberately building two SKUs with different trade-offs rather than one model for everything.
Grok 4.6's target date lands almost exactly 10 days after Kimi K3's full open-weight release rattled the industry. Chinese financial outlets reported the release wiped an estimated $314 billion off combined OpenAI/Anthropic valuation expectations — analyst estimates relayed through Chinese media, not independently confirmed, but directionally indicative of market sentiment.
How Grok stacks up against the field
| Model | Vendor | Parameters | Context | Pricing (input/output per 1M tokens) | Source |
|---|---|---|---|---|---|
| Grok 4.5 | xAI | Undisclosed | 500K | $2 / $6 | xAI official |
| Grok 4.6 (announced) | xAI | 1.5T | Undisclosed | Undisclosed | Musk's X post (unverified) |
| Kimi K3 | Moonshot AI | 2.8T (MoE, ~16/896 experts active) | 1M | $0.30 (cache hit) / $3 (cache miss) input, $15 output | Moonshot + Hugging Face |
| Claude Fable 5.1 (rumored) | Anthropic | Undisclosed | Undisclosed | Rumored unchanged from Fable 5 ($10 in / $50 out) | 36kr, WinCentral leaks — unconfirmed |
| GPT-5.6 Sol | OpenAI | Undisclosed | Undisclosed | Undisclosed | OpenAI official |
Grok 4.6 and Claude Fable 5.1 rows are pre-release claims, not verified benchmarks — useful for gauging release timing and positioning, not for head-to-head performance comparisons.
What to flag before you trust this timeline
- The only source is a tweet. No xAI blog post, model card, or product page confirms Grok 4.6's specs or date.
- "Musk time" has a track record. Treat "around August 7" as a target, not a guarantee.
- xAI is fighting safety controversies on a separate front. The July 2026 CSAM lawsuit implicitly concedes Grok can produce such content when safeguards are bypassed.
- Timing collides with an industry split on AI pacing. The same day Musk announced Grok 4.6/4.7, OpenAI and Anthropic endorsed "Pacing the Frontier" as companies. xAI is notably absent from that list.
Six steps to prepare before Grok 4.6 ships
With no official model card in hand, these six steps keep your evaluation rational through the August release crunch:
- Lock a Grok 4.5 baseline now: Run your production-relevant agent tasks through Grok Build or Cursor. Record token usage, latency, and success rates — that's your comparison anchor when 4.6 drops.
- Subscribe to official xAI channels: Follow xAI's X account and x.ai/blog. Musk's posts are early signals; the model card is the contract.
- Benchmark against shipped competitors: Use OpenRouter or direct APIs to run the same eval script against Kimi K3, Grok 4.5, and GPT-5.6 Sol. Multi-vendor baselines beat single-vendor hype.
- Run evals in an isolated environment: Agent benchmarks install dependencies, modify configs, and leave background processes. Use a disposable macOS instance — not your daily driver.
- Budget for pricing uncertainty: Grok 4.5 launched at $2/$6 per million tokens. Reserve 20–30% buffer for 4.6 pricing changes and keep vendor-switch clauses in any contract.
- Calendar two re-evaluation dates: August 7 (4.6 target) and August 21 (4.7 window). Re-run benchmarks against the official model card then — not against tweet specs now.
Hard numbers and sources you can cite
- Grok 4.5 Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%: From xAI's official Grok 4.5 model card. ~15,954 output tokens per SWE-Bench Pro task vs. Opus 4.8's 67,020.
- Kimi K3 Frontend Code Arena 1,679 points: First open-weight model to beat every closed model on that board. Artificial Analysis Intelligence Index rank ~3. Source: Moonshot AI official release.
- Grok 4.6 1.5T params, ~Aug 7 target: From Musk's July 28, 2026 X reply to @rauchg — not confirmed by xAI official announcement.
- Grok 4.7 2.1T params: Same source. Musk says "better than 4.6 in every way except slightly slower to serve."
Information current as of July 30, 2026. All Grok 4.6/4.7 details are vendor announcements or industry leaks — verify against official xAI channels before relying on any specific date, spec, or price.
Official xAI Grok 4.5 announcement:
Third-party coverage of Musk's July 28, 2026 X post (includes quoted text):
Tron Weekly — Elon Musk Targets August 7 For Grok 4.6
Moonshot AI Kimi K3 official release:
Hugging Face — moonshotai/Kimi-K3
Grok 4.6 release date FAQ
When exactly is Grok 4.6 coming out?
Musk said "around August 7, 2026" in an X post, but xAI has not officially confirmed a date. Treat it as a target that could shift.
What's the difference between Grok 4.6 and Grok 4.7?
Grok 4.6 is a 1.5T-parameter model focused on SFT/RL post-training improvements. Grok 4.7, expected a few weeks later, is a larger 2.1T model that Musk says outperforms 4.6 across the board except for serving speed, where it trades some latency for better token efficiency.
Will Grok 4.6 beat Kimi K3 or Claude Fable 5.1?
Too early to tell. Grok 4.6 has no published benchmarks yet, Kimi K3 already has verified third-party scores, and Claude Fable 5.1 hasn't even been officially confirmed by Anthropic.
How much will Grok 4.6 cost?
Unknown. Grok 4.5 launched at $2 per million input tokens and $6 per million output tokens — a reasonable reference point, but xAI hasn't disclosed Grok 4.6 pricing.
Where will I be able to use Grok 4.6?
Based on Grok 4.5's rollout, expect Grok Build, the xAI API, and the xAI console first, with third-party integrations following — but this isn't confirmed for 4.6 yet.
August 2026 is shaping up to be one of the densest model-release months on record — Grok 4.6, Grok 4.7, and a rumored Claude Fable 5.1 landing in the same window, weeks after Kimi K3's open-weight shock. That compresses the useful shelf life of any single flagship to a matter of weeks, which makes token efficiency and real-world task cost — not leaderboard rank alone — the more durable basis for a model choice. Running agent benchmarks and API comparisons at that pace pollutes a daily-driver machine fast: conflicting SDK versions, orphaned background agents, rewritten configs. KVMFLUX's cloud Mac mini rentals give you root access, SSH + VNC, and dedicated physical M4 hardware you can spin up for a day, benchmark, and tear down clean. If you're weighing rental periods, daily billing fits release-driven eval cycles best.
Benchmark Grok agent workflows on real Apple Silicon
Rent a Mac mini M4 by the day, run Grok 4.5 baselines and Cursor agent tests, and tear down before the August release wave hits your main machine.