A bigger model at the same price. Grok 4.7 was released on 21 September 2026, about six weeks after Grok 4.6, and SpaceXAI says it costs and runs the same: “Served at the same price and speed as Grok 4.6”. What changed is underneath: “Grok 4.7 uses a new, larger base model compared to Grok 4.6”.SpaceXAI, Introducing Grok 4.7, 21 Sep 2026, read at source 22 Sep 2026.
- Price, context window and reasoning settings are unchanged on the public API. xAI’s release notes describe both models in near-identical words.
- The model underneath is new: a larger base model and a longer reinforcement-learning run aimed at tasks that take hours.
- Every score is xAI’s own. The biggest reported gain is on multi-hour terminal work; none has been re-run independently.
- The comparison is not like for like: xAI tested 4.7 at its highest reasoning setting and 4.6 one step lower.
- Grok 4.6’s own scores moved between posts, so compare numbers from one table, never across two.
- The fast variant is Cursor and Grok Build only. It is not on the public API.
Every price here was correct when this site last checked it, on 22 Sep 2026. Prices change often and this site no longer updates them, so check the vendor’s own page before you rely on one: xAI pricing.
What stayed the same?
On the API, almost everything a developer sets or pays for. xAI’s release note for Grok 4.7 repeats the one it wrote for Grok 4.6 in August: a 500k context window, text and image in, text out, no text output limit, and the same four reasoning levels, low, medium, high (the default) and xhigh.xAI, Release Notes, entries for 21 September and 12 August 2026, read at source 22 Sep 2026: “It has a 500k context window, text and image inputs with text-only output, and no text output limit.”
The price is the same to the cent. Both entries give it in one sentence: $2 per million input tokens, $0.50 for cached input and $6 for output, doubling once a prompt passes 200k tokens. That is a snapshot from 22 September 2026: vendor prices change, so re-check the live page before budgeting.xAI, Release Notes, read at source 22 Sep 2026, in both the Grok 4.7 and Grok 4.6 entries: “Pricing is $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200k prompt tokens, and $4 / $1 / $12 above.”
Switching from 4.6 to 4.7 is a one-word change in the model name, not a budgeting decision. Whether it is worth doing depends on behaviour, which is where the price pages cannot help.
What changed under the hood?
The training, according to the launch post. 4.7 “was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete.” xAI adds: “The model is better at verifying its own work and managing longer context.” It also says it trained 4.7 to understand Grok Bot, the agent product xAI launched in August.SpaceXAI, Introducing Grok 4.7, 21 Sep 2026, read at source 22 Sep 2026: “We also trained Grok 4.7 to natively understand the Grok Bot harness”. xAI, Release Notes, 11 August 2026 entry: “Durable AI teammates that work on a persistent cloud computer, with messaging, approvals, connectors, and routines.”
That continues the direction of the last release. Grok 4.6 was also pitched at long-running agents, and xAI said then that it had started to see the model checking its own work (see Grok 4.6 vs 4.5).
xAI’s models page gives 4.7 a knowledge cut-off of May 2026 and now tells developers to use it for everything except images, video and voice.xAI, Models, last updated 21 Sep 2026, read at source 22 Sep 2026: “The knowledge cut-off date of Grok 4.7 is May 2026.” and “For everything else, including code, use Grok 4.7. It is the most capable model we’ve built.”
On safety, the post claims “an entirely new safeguard stack”, a score of 62.4% on LatchBio’s biosafety benchmark, and that on xAI’s own cyber benchmark the model let through “only 3.3% of risky dual-use prompts”. Both figures come from xAI, and the second from a benchmark it built.SpaceXAI, Introducing Grok 4.7, 21 Sep 2026, read at source 22 Sep 2026: “topping LatchBio’s biosafety benchmark at 62.4%”.
A new, larger base model is a bigger change than the version number suggests. Same price does not mean same behaviour: prompts tuned on 4.6 may need re-testing.
What do the benchmarks say?
| Benchmark (what it tests, per xAI) | Grok 4.7 | Grok 4.6 |
|---|---|---|
| Terminal-Bench 4.0 (multi-hour terminal work) | 38.0% | 20.3% |
| EEBench (electrical engineering) | 64.0% | 53.0% |
| HealthBench Professional (clinical reasoning) | 56.7% | 48.5% |
| CursorBench 4.0 (software engineering) | 46.3% | 40.4% |
| DeepSWE v1.1 (software engineering) | 71.0% (high effort) | 65.2% |
| Harvey Legal Agent Benchmark (legal work) | 19.6% | 15.8% |
| AA Briefcase v1.1 (multi-hour office work) | 1,657 | 1,546 |
| GDPval (Elo score) | 1695 | 1605 |
SpaceXAI, Introducing Grok 4.7, 21 Sep 2026, read at source 22 Sep 2026. xAI’s own table, which describes itself as “Token prices and benchmark scores for Grok 4.7, Grok 4.6, GPT-5.6 Sol, and Fable 5.1.” No score here was reproduced independently.
In xAI’s own figures, 4.7 is ahead of 4.6 on all eight. The gap is widest on Terminal-Bench 4.0, where the score nearly doubles, and narrowest on the legal benchmark and the two software-engineering tests.
Two things make these rows less clean than they look. First, the settings differ: the table puts 4.7 at xHigh against 4.6 at High, yet xAI’s own release note says 4.6 also supports xhigh. The DeepSWE figure for 4.7 is marked as a high-effort score, so that one row is closer to like for like.SpaceXAI, Introducing Grok 4.7, read at source 22 Sep 2026: “An asterisk on Grok 4.7 DeepSWE marks a high-effort score.” xAI, Release Notes, 12 August 2026 entry for Grok 4.6: “Reasoning effort supports low, medium, high (default), and xhigh.” Second, the 4.7 post carries no note on where the scores came from. The 4.6 post did.SpaceXAI, Introducing Grok 4.6, 12 Aug 2026, read at source 22 Sep 2026: “Third-party model scores are the best of self-reported or publicly available results.”
The direction is consistent, and the biggest claimed gains line up with what xAI says it trained for: long, many-step work. How big the gains are, at matched settings, nobody outside xAI has measured yet.
Why 4.6’s own scores changed
Put the August post next to the September one and Grok 4.6 scores differently in each. Some of that is new test versions. In August xAI reported 69.9% on CursorBench v3.2; in September it reported 40.4% on CursorBench 4.0, a newer version of the test.SpaceXAI, Introducing Grok 4.6, 12 Aug 2026, and Introducing Grok 4.7, 21 Sep 2026, both read at source 22 Sep 2026: table rows “CursorBench v3.2 69.9%” (August) and “CursorBench 4.0 46.3% 40.4%” (September, Grok 4.7 then Grok 4.6).
One row changed with the version number unchanged. DeepSWE v1.1 is named in both tables, both times for Grok 4.6 at High, and the score is 65.9% in August and 65.2% in September. Neither post explains the difference.SpaceXAI, Introducing Grok 4.6 (table row “DeepSWE v1.1 65.9%”) and Introducing Grok 4.7 (row “DeepSWE v1.1 71.0%* 65.2%”), both read at source 22 Sep 2026. It is a small gap, but it means a score is a result from one run of one harness on one day, not a fixed property of the model.
Our Grok 4.6 vs 4.5 page quotes the August table and stays accurate to it. Do not mix its numbers with the ones on this page.
Only compare two models inside one table from one date. A score lifted from an older announcement may have been measured on a different version of the test.
The fast variant, and where 4.7 runs
The launch post says: “We also serve a fast variant with twice the output speed at twice the price.” The developer docs narrow where you can get it: only in Cursor and Grok Build, not through the public API.xAI, Release Notes, 21 September 2026, read at source 22 Sep 2026: “Grok 4.7 Fast, the same model at twice the token rates, is available only through Cursor and Grok Build, not on the public xAI API.” Grok 4.6 also had a fast tier at double the price; its post did not say where it was sold.SpaceXAI, Introducing Grok 4.6, read at source 22 Sep 2026: “Additionally, there is a fast variant which is twice the price.”
Standard 4.7 runs on the xAI API, in Cursor, as the default model in Grok Build, and through OpenRouter, Vercel and Cloudflare.xAI, Grok 4.7, read at source 22 Sep 2026: “OpenRouter, Vercel, and Cloudflare”. xAI’s US-only endpoint, at a 10% premium, is not a 4.7 difference: it serves both models.xAI, Pricing, read at source 22 Sep 2026: “a 10% premium”, models “Currently grok-4.7 and grok-4.6 only”.
One API detail appears in the 4.7 release note and not the 4.6 one: on the Responses API, 4.7 “always returns reasoning.encrypted_content”, so client code that stores or filters responses may see an extra field.xAI, Release Notes, 21 September 2026, read at source 22 Sep 2026: “On the Responses API, grok-4.7 always returns reasoning.encrypted_content”.
If you call the public API, the fast variant is not an option for you. The standard 4.7 is what you get, at the 4.6 price.
Should you switch?
At the same price, the only cost of trying 4.7 is re-testing. Run the prompts you already rely on through both models and compare the output, rather than trusting either table. Grok 4.6 is still on xAI’s price list as of 22 September 2026, so there is no forced move yet.xAI, Pricing, read at source 22 Sep 2026: grok-4.6 listed at “500k” context, the same rates as grok-4.7. If you need results that do not shift under you, xAI says a model name with a date on the end “refers directly to a specific model release”.xAI, Models, read at source 22 Sep 2026.
For how Grok’s price sits against other vendors, see LLM API prices compared and Gemini vs Grok API prices. The company signing these posts is SpaceXAI, the renamed xAI: who owns Grok explains the SpaceX acquisition.
Same price makes this an easy trial and a hard claim to judge. Your own prompts are the only benchmark that answers your question.
What this page could not verify
- Every benchmark and safety number. All are xAI’s own; none was re-run here or found independently reproduced.
- “Twice as fast, at half the price of comparable models.” The launch post’s headline line. Which models count as comparable is xAI’s choice.
- Grok 4.6’s knowledge cut-off. Its developer page now redirects to the Grok 4.7 page, so there is nothing to compare against.
- Why Grok 4.6’s DeepSWE score changed. Neither post explains it.
- When Grok 4.6 will be retired. No date was published as of 22 September 2026.
SpaceXAI, Introducing Grok 4.7, 21 September 2026 — Introducing Grok 4.6, 12 August 2026. xAI developer docs: Release Notes, Models, Pricing and Grok 4.7. All read at source on 22 September 2026.
Prices are a snapshot and vendor benchmarks are marketing until someone else runs them.