ONLINEAGENT_OPS 2026.Q3 HOME ARTICLES CRAFT RECORD BLOG MAP HUBS FAQ SEARCH
HOMEARTICLESGROK 4.6 VS GROK 4.5: WHAT ACTUALLY CHANGED
ARTICLES · COMPARISON

Grok 4.6 vs Grok 4.5: What Actually Changed

Four weeks apart, same price. The published benchmark gaps, the one behaviour change worth noting, and why Grok 4.3 is not in this comparison.

READ3 min
WORDS646
SECTIONS5
TYPECOMPARISON
CHECKED16 SEP 26

Same price, better at long jobs. Grok 4.6 landed on 12 August 2026, four weeks after 4.5, and xAI describes it as an upgrade aimed at agents that keep working: “Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.”SpaceXAI, Introducing Grok 4.6, 12 Aug 2026, read at source 16 Sep 2026.

◈ NEWER MODEL OUT

Grok 4.7 was released on 21 September 2026 at the same price; what changed is on Grok 4.7 vs 4.6. This page stays as the record of the 4.5 to 4.6 step.

TL;DR — THE SHORT VERSION

An incremental release that moved the agentic scores and left the price alone.

  • Prices are identical: $2 per million input tokens, $6 per million output, on both.SpaceXAI, Introducing Grok 4.6, 12 Aug 2026, and Introducing Grok 4.5, 16 Jul 2026: “$2 per million input tokens and $6 per million output tokens” on each.
  • The gap is largest on agent benchmarks, not on general intelligence scores.
  • xAI claims parity with GPT-5.6 Sol on one composite index.
  • 4.6 reports self-checking — xAI says it saw the model verifying its own work on long tasks.
  • Grok 4.3 is not covered here. Its announcement page does not resolve.
◈ PRICES ON THIS PAGE

Every price here was correct when this site last checked it, on 16 Sep 2026. Prices change often and this site no longer updates them, so check the vendor’s own page before you rely on one: xAI pricing.

What actually changed?

xAI’s published comparison, Grok 4.6 High against Grok 4.5 High, read at source 16 Sep 2026
BenchmarkGrok 4.6Grok 4.5
AA Intelligence Index6156
DeepSWE v1.165.9%54%
APEX-Agents57.5%47.1%
Terminal-Bench v3.026%15.7%
CursorBench v3.269.9%66.7%
FrontierCode v1.1 (Extended)61.3%56.6%

SpaceXAI, Introducing Grok 4.6, 12 Aug 2026. The table is xAI’s own; it notes “Third-party model scores are the best of self-reported or publicly available results.” No score here was reproduced independently.

GROK 4.6 AGAINST 4.5 · xAI’S OWN SCORES
The percentage rows of the table, drawn to scale. The first three pairs are the task-sticking benchmarks where the jumps are.
Grok 4.6Grok 4.5
DeepSWE v1.1 · Grok 4.665.9%
DeepSWE v1.1 · Grok 4.554%
APEX-Agents · Grok 4.657.5%
APEX-Agents · Grok 4.547.1%
Terminal-Bench v3.0 · Grok 4.626%
Terminal-Bench v3.0 · Grok 4.515.7%
FrontierCode v1.1 (Extended) · Grok 4.661.3%
FrontierCode v1.1 (Extended) · Grok 4.556.6%
CursorBench v3.2 · Grok 4.669.9%
CursorBench v3.2 · Grok 4.566.7%
SpaceXAI, Introducing Grok 4.6, 12 Aug 2026, read at source 16 Sep 2026. xAI’s own comparison, Grok 4.6 High against Grok 4.5 High; it notes “Third-party model scores are the best of self-reported or publicly available results.” No score here was reproduced independently.
TAKEAWAY

The jumps are in the columns that measure sticking with a task — DeepSWE, APEX-Agents, Terminal-Bench. The general-knowledge index moved five points. That is the shape of an agent release.

The claim worth reading twice

xAI says of 4.6: “It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.” A composite of nine is a useful summary and a poor answer to whether it will do your job. Both models sit at 61 in xAI’s own table.

This is a vendor comparing itself to a competitor using a third party’s index. The index is real; the selection of it is not neutral.

The behavioural difference

The most interesting line in the announcement is not a number: “On longer trajectories, we also started to see more self-testing and verification, with the model checking its own work before moving on.” xAI also says 4.6 “produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5”.

How it got there is described plainly — a longer supplemental training run, then Grok 4.5 used “to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains”. The older model trained the newer one.

Speed and cost

4.5 was the speed story: xAI says it “is served at fast-model speeds of 80 TPS” and that it “resolves tasks with 15,954 output tokens on average” on SWE Bench Pro. 4.6 keeps the price and adds a faster tier: “there is a fast variant which is twice the price.”SpaceXAI, Introducing Grok 4.5, 16 Jul 2026, and Introducing Grok 4.6, 12 Aug 2026, read at source 16 Sep 2026: “Grok 4.5 resolves tasks with 15,954 output tokens on average”.

Both were trained on xAI’s own cluster — 4.5 “across tens of thousands of NVIDIA GB300 GPUs”. The machine behind that is Colossus, and the company signing these posts is now SpaceXAI: the company that launched Grok, renamed after SpaceX acquired it.

What this page could not verify

  • Grok 4.3. x.ai/news/grok-4-3 returned 404 on 16 Sep 2026, so there is nothing to quote — although the model is still sold: xAI’s pricing page lists it at $1.25 and $2.50 per million tokens with a 1M context, read at source the same day.xAI, Pricing, read at source 16 Sep 2026: grok-4.3 “$1.25” input and “$2.50” output per 1M tokens.
  • Every benchmark number. All are xAI’s, none re-run here.
  • Whether 4.5 is still available. The announcements do not say it was retired.
SOURCES

SpaceXAI, Introducing Grok 4.6, 12 August 2026 · Introducing Grok 4.5, 16 July 2026. Both read at source on 16 September 2026.

Vendor benchmarks are marketing until someone else runs them.

ABOUTMETHODVERIFYPRIVACYCONTACTINDEXAI PROMPT GENEER · EVERY ARTICLE CARRIES ITS OWN CHECKED DATE