
xAI released Grok 4.7 on September 21, 2026 (UTC-4), after at least five delays. The new flagship weighs in at 2.1 trillion parameters — up from 1.5 trillion in Grok 4.6 — yet keeps the same $2/$6 per million token pricing. It's available now on the Grok app, Cursor, Grok Build, the xAI API, and GitHub Copilot. Elon Musk said Grok 4.8 is already trained.
What shipped
Grok 4.7 replaces Grok 4.6 as xAI's top model. The parameter jump from 1.5T to 2.1T is substantial — a 40% increase — and xAI claims "significant improvements" at the same price and speed. The model includes a new safety stack and grants selected cybersecurity partners access for defensive research. Knowledge cutoff is May 2026.
The distribution strategy is notable for launch day breadth. Grok 4.7 isn't just on xAI's own platforms — it's live in Cursor, GitHub Copilot, and the xAI API from day one. That's a direct play for developer market share, where Claude Code and GPT-6 Astra currently dominate.
The benchmark picture — and the fight
Here's where it gets interesting. xAI's own numbers and independent evaluators don't agree on coding performance.
| Benchmark | Grok 4.7 | Claude Fable 5.1 | GPT-6 Astra | Source |
|---|---|---|---|---|
| GDPval | 1695 (2nd) | 1735 (1st) | — | xAI / AA |
| AA-Briefcase | 1657 (2nd) | 1678 (1st) | — | xAI / AA |
| EEBench (electrical eng.) | 2nd | — | 1st | xAI |
| CursorBench 4.0 | 46.3 | — | — | Cursor |
| AA Composite Intelligence | 46 | 53 | 53 | Artificial Analysis |
| Coding (xAI claim) | 38% | — | — | xAI |
| Coding (independent) | 26% | — | — | Artificial Analysis |
The 12-point gap between xAI's 38% coding claim and Artificial Analysis's 26% independent reading is the headline controversy. That's not a rounding difference — it's a 46% relative gap. xAI attributes the difference to methodology; Artificial Analysis stands by its evaluation pipeline. Either way, Grok 4.7 sits firmly in second place on most benchmarks, behind both Claude Fable 5.1 and GPT-6 Astra on the AA Composite Intelligence index (46 vs. 53).
The AA Composite score is particularly telling: Grok 4.7 gained only 2 points over Grok 4.6, while the frontier models sit at 53. That's a 7-point gap to the leaders, and the rate of improvement from 4.6 to 4.7 doesn't suggest xAI is closing it quickly.
Why it matters
Grok 4.7's real story isn't the benchmarks — it's the pricing. At $2/$6 per million tokens, Grok 4.7 costs roughly half of Claude Opus 5 ($5/$25) and a fifth of GPT-6 Astra's standard pricing ($10/$50). If Grok 4.7 delivers 85-90% of frontier performance at 20-50% of the cost, that changes the economics for price-sensitive workloads — code generation, data processing, agentic tasks where volume matters more than peak intelligence.
The 2.1 trillion parameter count at $2/$6 is also a statement about xAI's cost structure. Training and serving a 2.1T model cheaply enough to charge $2/$6 implies either serious infrastructure efficiency or aggressive subsidization. xAI's access to SpaceX/Memphis data center capacity and Colossus GPU clusters likely plays a role, but the unit economics here are worth watching. If xAI can sustain this pricing profitably, it puts pressure on every other provider to justify their premiums.
The five delays matter too. Grok 4.7 was supposed to ship earlier this year, and the repeated slips suggest xAI's training pipeline hit real obstacles — whether data quality, alignment issues, or infrastructure bottlenecks. The fact that Musk immediately announced Grok 4.8 is "already trained" could be read as confidence, or as an attempt to distract from 4.7's underwhelming benchmark gains.
The critical lens
Let's be honest about what Grok 4.7 is and isn't. It's a solid second-tier model with aggressive pricing and broad distribution. It is not a frontier model. The 2-point gain on AA Composite, the second-place finishes across every major benchmark, and the 12-point coding discrepancy all point to the same conclusion: xAI is competing on price and distribution, not on raw intelligence.
The benchmark discrepancy deserves scrutiny beyond just "methodology differences." When a vendor's self-reported coding score is 46% higher than an independent evaluator's, users should ask what task selection and prompting strategies xAI used. This isn't unique to xAI — every vendor optimizes for benchmarks — but the size of the gap here is unusual and warrants independent verification before enterprises commit to volume pricing.
Musk's announcement that Grok 4.8 is already trained raises its own questions. If 4.8 is ready, why ship 4.7 at all? The answer may be that 4.8 isn't actually ready for production — it may be a research checkpoint, or it may have alignment issues that need work. Using "4.8 is trained" as a marketing hook while shipping 4.7 is classic Musk: create forward momentum to offset underwhelming current results.
What to watch
The first real test is whether Cursor and GitHub Copilot users actually switch. Grok 4.7's day-one integration with those platforms means it gets distribution without users having to seek it out. Watch Cursor's model usage stats over the next 2-4 weeks — if Grok 4.7 captures more than 10% of coding sessions, the pricing strategy is working. If it stays under 5%, the benchmark gap matters more than the price gap.
Watch for Artificial Analysis and other independent evaluators to publish full Grok 4.7 results in the coming days. The 12-point coding discrepancy will either be resolved or widened — and either outcome matters for enterprise trust.
And watch Grok 4.8. If it ships within 2-3 months with genuine frontier-level benchmarks, xAI's rapid iteration cycle becomes a real competitive threat. If it slips again or ships with similar second-tier performance, the "4.8 is trained" announcement will look like what it probably is: a distraction.
No comments yet