Skip to content
news19 min read

Grok 4.6 Ties GPT-5.6 Sol on the AA Intelligence Index — and Ranks Below Grok 4.5 on LMArena

Grok 4.6 ties GPT-5.6 Sol on AA's Intelligence Index — yet ranks below Grok 4.5 on LMArena, and its cost lead is now 1.7%. Four instruments, four answers.

Author
Anthony M.
19 min readVerified August 28, 2026Tested hands-on
Grok 4.6 scores 60.92 on the Artificial Analysis Intelligence Index v4.1.1 against 60.93 for GPT-5.6 Sol, measured at high effort against max effort
SpaceXAI released Grok 4.6 on August 12, 2026. On the Artificial Analysis Intelligence Index v4.1.1 it reads level with GPT-5.6 Sol — at a different reasoning effort.

Update, August 28, 2026. Two things on this page have been corrected. First, the cost gap named in the original headline no longer holds: Artificial Analysis has re-measured cost per task, and Grok 4.6 at high now reads $0.9372 per index task against $0.9530 for GPT-5.6 Sol at max — 1.7 percent apart, where the August 22, 2026 reading showed $0.837 against $1.231. The index scores themselves have not moved: 60.92 against 60.93 on both readings. Second, this article originally stated that Artificial Analysis published no score for Grok 4.6 at xhigh. That was wrong. The evaluator scores all four reasoning efforts on the model release page, and the xhigh result is lower than high on the Intelligence Index, not higher. Both corrections are carried through the text, tables and FAQ below.

The Gist: SpaceXAI released Grok 4.6 on August 12, 2026, and on one independent instrument it now sits level with a top OpenAI model. On the Artificial Analysis Intelligence Index v4.1.1, in readings taken on August 22 and again on August 28, 2026, Grok 4.6 scores 60.92 against 60.93 for GPT-5.6 Sol — a gap of seven thousandths of a point — while the cost gap between them has closed: at the August 28, 2026 reading Grok 4.6 runs the index for $0.9372 per task against $0.9530 for Sol, a difference of 1.7 percent, where the August 22 reading showed $0.837 against $1.231. On that same index it still trails Claude Opus 5 at 63.05 and Claude Fable 5 at 62.07. The configuration matters and is not symmetrical: Grok 4.6 is measured at high, which SpaceXAI's own harness documents as its default effort, while GPT-5.6 Sol and Claude Opus 5 are measured at max and Claude Fable 5 under the configuration Artificial Analysis labels with fallback. Artificial Analysis scores all four of Grok 4.6's reasoning efforts, and xhigh reads lower than high on the index — 60.01 against 60.92 — for about 31 percent more per task. Three other instruments read differently: on LMArena, read August 22, 2026, the model ranks 46th with 3,473 votes and a pre-release flag, below its own predecessor Grok 4.5 at 35th; on the Artificial Analysis Coding Agent Index v1.4 it is not measured at all; and on SpaceXAI's own comparison table it posts the leading figure in 3 of 10 rows.

Key Takeaways

  • The tie is real, and it is one number on one index. Artificial Analysis Intelligence Index v4.1.1, read August 22, 2026: Grok 4.6 at high scores 60.92, GPT-5.6 Sol at max scores 60.93. That is a tie in any practical sense — and it is a composite of nine evaluations, not a verdict on everything a model does.
  • The price argument was strong on August 22 and is thin now. On August 22, 2026 Artificial Analysis measured $0.837 per index task for Grok 4.6 against $1.231 for GPT-5.6 Sol. Six days later that gap had almost closed: the August 28 reading is $0.9372 against $0.9530, a difference of 1.7 percent, with $2.337 for Claude Opus 5 and $3.140 for Claude Fable 5 unchanged across both. Grok 4.6 remains the cheaper of the two, and it was never the cheapest thing on the board: GLM-5.3 at max runs the same index for $0.683.
  • The effort levels are not matched, and nobody should pretend otherwise. SpaceXAI documents four reasoning efforts for Grok 4.6 — low, medium, high, and xhigh — and Cursor's documentation names high as the default. Artificial Analysis does publish xhigh figures — the model scores 60.01 there, below the 60.92 it posts at high. Anyone writing "at equal settings" about this comparison is still describing a run nobody made: Sol was measured at max and Grok 4.6 at high.
  • On human preference it ranks below the model it replaces. LMArena's text leaderboard, read August 22, 2026, places grok-4.6-high at rank 46 with an Elo of 1461.16, against rank 35 and 1470.29 for Grok 4.5. The reserve is large: 3,473 votes, a pre_release flag, and a rank interval spanning 25 to 69.
  • Its own launch chart leaves out the model currently leading the index it cites. SpaceXAI's comparison table sets Grok 4.6 against Grok 4.5, GPT-5.6 Sol, and Claude Fable 5. Claude Opus 5, which tops the Artificial Analysis Intelligence Index at 63.05 on the same date, does not appear in it.

What SpaceXAI Actually Shipped on August 12

SpaceXAI announced Grok 4.6 on August 12, 2026, positioning it as a successor to Grok 4.5 aimed at long-running agents and, in the company's words, "more ambitious interactive and visual work." The company is the entity formerly known as xAI, which completed its rebrand earlier in 2026; we covered that change in xAI Is Officially SpaceXAI Now, and the model line kept the Grok name.

The hard specifications come from the vendor's model page. According to SpaceXAI's documentation for grok-4.6, the model exposes a 500,000-token context window, accepts text and image input and returns text, and supports function calling, structured outputs, and configurable reasoning. Its knowledge cutoff is February 1, 2026. It is served from two regions, us-east-1 and us-west-2, with published rate limits of 150 requests per second and 50 million tokens per minute.

List pricing, from SpaceXAI's API pricing page, has two tiers. Short context bills $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens. Long context — which the page defines as any request whose prompt reaches 200,000 tokens — bills $4.00, $1.00, and $12.00 respectively.

grok-4.6 list priceInputCached inputOutput
Short context (below 200k tokens)$2.00 / 1M$0.50 / 1M$6.00 / 1M
Long context (prompt ≥ 200k tokens)$4.00 / 1M$1.00 / 1M$12.00 / 1M

Source: SpaceXAI API pricing page, read August 22, 2026. The long-context threshold is the instrument here — see the section on where the sources disagree.

One detail in that table is easy to miss and expensive to learn late: the pricing page states that models with long-context pricing bill the long-context rates for all tokens in a request once its prompt reaches the threshold. Crossing 200,000 tokens does not surcharge the overflow. It reprices the entire call.

Worth noting against its predecessor: input and output rates are identical to Grok 4.5 at $2.00 and $6.00, but the cached-input rate moved from $0.30 to $0.50 per million tokens. For a cache-heavy agent loop, the newer model is the more expensive one on that line.

The top four models on the Artificial Analysis Intelligence Index v4.1.1 read on August 22, 2026, each with the reasoning effort it was measured at: Claude Opus 5 at max 63.05, Claude Fable 5 with fallback 62.07, GPT-5.6 Sol at max 60.93 and Grok 4.6 at high 60.92
The top of the Artificial Analysis Intelligence Index v4.1.1, read August 22, 2026. Grok 4.6 ties GPT-5.6 Sol on score and undercuts it on cost per task.

The Index Where Grok 4.6 Ties a Frontier OpenAI Model

The short answer: on the Artificial Analysis Intelligence Index v4.1.1, read on August 22, 2026, Grok 4.6 at high effort scores 60.92 and GPT-5.6 Sol at max scores 60.93. The two are level to within seven thousandths of a point. What separated them on cost has since narrowed: at the August 28, 2026 reading Grok 4.6 runs the same battery for 1.7 percent less, where the August 22 reading showed about a third less.

The index is a composite. Artificial Analysis defines v4.1.1 as nine evaluations combined: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. The evaluator states that results are measured independently on dedicated hardware. Here is the top of that board on the day we read it.

That version number is not decoration, because the instrument itself moves. Artificial Analysis records what changed in v4.1.1: 𝜏³-Banking was updated to v1.0.1, and the graders for Humanity’s Last Exam, AA-LCR and AA-Omniscience were upgraded to GPT-5.6 Luna (medium). The practical consequence is worth stating plainly: a score read under an earlier version is not comparable to one read under v4.1.1. When a model’s number moves between two readings taken months apart, the model may not have changed at all — the measurement did. Every Artificial Analysis figure on this page comes from one reading of v4.1.1, taken on August 22, 2026, which is why no score here is presented as a gain or a loss over time.

Model (effort as measured)AA Intelligence Index v4.1.1Cost per index task
Claude Opus 5 (max)63.05$2.337
Claude Fable 5 (with fallback)62.07$3.140
GPT-5.6 Sol (max)60.93$0.9530
Grok 4.6 (high)60.92$0.9372
Kimi K3 (max)59.70$0.837
GLM-5.3 (max)59.51$0.683

Source: Artificial Analysis, Intelligence Index v4.1.1 and Cost per Task datasets. Index scores were read August 22, 2026 and are unchanged on re-reading August 28, 2026; cost per task is the August 28, 2026 reading, which moved for Grok 4.6 and GPT-5.6 Sol and not for the other rows. Effort levels are as the evaluator ran them and are not uniform across rows.

Two things in that table deserve to be said out loud rather than left for the reader to infer. First, the models above Grok 4.6 are still above it: Claude Opus 5 leads by 2.13 points and Claude Fable 5 by 1.15 points on this index. Second, the cost column does not make Grok 4.6 the budget option in any absolute sense. Kimi K3 now runs the same battery for $0.837 — about 11 percent less than Grok 4.6 — while scoring 1.2 points lower, and GLM-5.3 runs it for $0.683, around 27 percent less, while scoring 1.4 points lower. The honest framing is narrow: among the four models at the top of this index Grok 4.6 is still the cheapest per task, but its lead over GPT-5.6 Sol is 1.7 percent, not a third.

Artificial Analysis also measures a median output speed of 68.7 tokens per second for Grok 4.6 at high, which places it between GPT-5.6 Sol at 69.98 and Claude Opus 5 at 58.39 on the same reading. This is not a fast model relative to the wider field — Gemini 3.7 Flash at high measures 389.5 on the same instrument.

Grok 4.6's four reasoning efforts drawn as a staircase — low, medium, high and xhigh — with high, its default, crowned as the highest of the four on the Artificial Analysis Intelligence Index, and xhigh, the top setting, standing taller but marked below high on that index
Grok 4.6 ships four reasoning efforts and Artificial Analysis scores all four. Its default, high, is the highest of the four on the Intelligence Index; xhigh, the top setting, sits below it.

The Asterisk That Decides How Much the Tie Is Worth

The short answer: the two tied entries were not run at the same reasoning effort. Grok 4.6 is measured at high, its default; GPT-5.6 Sol is measured at max. Artificial Analysis does score Grok 4.6 at xhigh, and the result runs against intuition: 60.01 on the index, below the 60.92 it posts at high, for about 31 percent more per task.

SpaceXAI documents four reasoning efforts for Grok 4.6 — low, medium, high, and xhigh — in its Amazon Bedrock announcement, and Cursor's model documentation states plainly that high is the default. So the index is comparing a model at its shipped default against competitors measured under settings other than its own — max for GPT-5.6 Sol and Claude Opus 5, and for Claude Fable 5 the configuration Artificial Analysis labels with fallback.

That asymmetry is real, but it does not cut the way most coverage assumes. Artificial Analysis scores all four of Grok 4.6's reasoning efforts on the model release page, and pushing the model past its default does not raise its index score — it lowers it. Here are the four, read August 28, 2026.

Reasoning effortAA Intelligence Index v4.1.1AA-OmniscienceAA-Briefcase EloCost per index task
low51.6825.901309.99$0.2545
medium59.0128.001521.82$0.7829
high (default)60.9230.481562.08$0.9372
xhigh60.0129.321587.28$1.2304

Source: Artificial Analysis, Grok 4.6 release page, read August 28, 2026. These are single readings and the evaluator publishes no error bars for them.

Two of the three quality instruments read lower at xhigh than at high: the Intelligence Index drops 0.91 points and AA-Omniscience drops 1.17, while AA-Briefcase Elo rises 25.20 points. Cost rises about 31 percent either way. We are not claiming high is the better setting everywhere — these are single readings, the differences are small, and no error bars are published. The narrower claim is the one the numbers support: the assumption that the top effort setting is the one to reach for when accuracy matters is not backed by this measurement.

Higher effort does still cost more money — 31 percent more per task for Grok 4.6 between high and xhigh. Between the two tied entries, though, that consequence has all but vanished: $0.9372 at high for Grok 4.6 against $0.9530 at max for Sol. The asymmetry of settings remains, and it still means the two were not run under comparable conditions. What no longer follows from it is a price argument.

The only defensible statement is the narrow one: at the configurations Artificial Analysis actually ran, the two models score the same on this index. What that costs has moved: the August 22, 2026 reading put Grok 4.6 about a third below Sol, and the August 28 reading puts the two 1.7 percent apart. "Matches at equal settings" is still a claim about a run nobody made. "Beats it when you push it" is now wrong for a sharper reason — pushed to xhigh, the model scores lower.

The Board Where It Ranks Below Its Own Predecessor

The short answer: on LMArena's text leaderboard, read August 22, 2026, grok-4.6-high sits at rank 46 with an Elo of 1461.16, while Grok 4.5 sits at rank 35 with 1470.29. The newer model is placed below the older one — on a provisional entry with 3,473 votes and a pre-release flag.

LMArena measures something different from Artificial Analysis: blind human preference between two anonymous model responses, aggregated into an Elo rating. It is not a capability battery, and a score here does not transfer to the index above or the other way around.

EntryRankEloConfidence intervalVotes
claude-fable-511507.771502.58 – 1512.9524,331
claude-opus-5-high71493±527,610
gpt-5.6-sol-xhigh161482.261476.86 – 1487.6619,541
grok-4.5351470.291465.13 – 1475.4522,030
grok-4.6-high461461.161451.07 – 1471.263,473

Source: LMArena text leaderboard payload, read August 22, 2026. The grok-4.6-high entry carries a pre_release flag and displays as preliminary.

The reserve on that row is not a formality, and we would rather state it than bury it. LMArena marks the entry pre_release and displays it as preliminary. Its rank interval runs from 25 to 69 — a 44-place span. And its confidence interval, 1451.07 to 1471.26, overlaps Grok 4.5's interval of 1465.13 to 1475.45 across the range 1465.13 to 1471.26. In plain terms: the board currently places Grok 4.6 below Grok 4.5, but with this many votes it does not cleanly separate the two. That ordering can move.

What is harder to wave away is the company it keeps. Five SpaceXAI entries rank above grok-4.6-high on the same board — grok-4.20-beta1 at 25, grok-4.20-beta-0309-reasoning at 33, grok-4.5 at 35, grok-4.20-multi-agent-beta-0309 at 36, and grok-4.1-thinking at 42. For a flagship release, being placed behind five of your own earlier entries on a human-preference board is a signal about conversational preference, even a provisional one — and it is not something the launch chart shows.

The Board Where It Is Not Measured at All

The short answer: on the Artificial Analysis Coding Agent Index v1.4, read August 22, 2026, there is no Grok 4.6 entry. Across the 55 harness-and-model pairs the board carries, the most recent SpaceXAI pairing is "Grok Build - Grok 4.5 (high)."

This matters because SpaceXAI presents Grok 4.6 as a model for long-running agents and coding, and the Coding Agent Index is the instrument built for exactly that. Version 1.4 combines three benchmarks — DeepSWE, SWE-Atlas-QnA, and Terminal-Bench v2.1 — and, crucially, it scores pairs: a harness plus a model plus an effort level, not a model in the abstract. "Claude Code - Opus 5 (xhigh)" and "Codex - GPT-5.6 Sol (max)" are separate entries because the harness is part of what is being measured.

One caveat on our own method, since it applies to anyone checking this: the structured data published in that page's markup exposes only nine display rows. Concluding "not ranked" from those nine would be wrong. We read the full board payload, counted 55 pairs, and found no Grok 4.6 pairing at any effort. Grok Build with Grok 4.5 at high scores 0.641 there at a mean $2.44 per task, against 0.681 for Claude Code with Opus 5 at xhigh at $8.17 per task. We covered the harness itself in SpaceXAI Open-Sources Grok Build.

An absence is not a result. It does not mean Grok 4.6 would score badly on that index; it means that on August 22, 2026 the independent agentic-coding board has not yet run it, so anyone citing Grok 4.6's agentic-coding standing is citing the vendor, not an evaluator.

What the Vendor's Own Table Says — and What It Leaves Out

The short answer: on SpaceXAI's own comparison table, Grok 4.6 posts the leading figure in 3 of 10 rows. These are vendor-published numbers, and the company states that competitor figures are self-reported or drawn from public results.

The table below is reproduced from SpaceXAI's launch post, which carries the note: "Third-party model scores are the best of self-reported or publicly available results." That footnote is the instrument. These are not independent measurements, and we present them as the company's claims.

Evaluation (vendor-published)Grok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Source: SpaceXAI launch post, read August 22, 2026. Bold marks the leading figure as the vendor marks it. Competitor figures are self-reported or publicly available results per the vendor's own note.

Three observations, all factual. Grok 4.6 leads 3 of the 10 rows — GDPVal-AA v2, AA-Briefcase, and Harvey LAB. Claude Fable 5 leads five of them and GPT-5.6 Sol leads two, on a chart the vendor chose. The gap is widest where the task is agentic and terminal-shaped: 26 percent against 34.6 percent on Terminal-Bench v3.0, and 65.9 against 73 on DeepSWE v1.1, both against GPT-5.6 Sol.

One reconciliation, because the two tables on this page disagree at a glance. The vendor chart rounds to whole numbers. It prints 61 for both Grok 4.6 and GPT-5.6 Sol and 62 for Claude Fable 5, where the Artificial Analysis figures for that same index are 60.92, 60.93 and 62.07. Rounding is why the vendor’s chart can show a flat tie at 61. It cuts the other way too: on a rounded board, two models separated by several tenths of a point can print the same integer and look level when they are not. That is the reason every index figure in this article is given to two decimals.

The improvement over its own predecessor is the cleanest story in the table: as the vendor presents it, Grok 4.6 is ahead of Grok 4.5 in all ten rows, by 227 points on GDPVal-AA v2 and by more than 10 points on Terminal-Bench v3.0 and DeepSWE v1.1. That comparison is the vendor’s own, drawn on the vendor’s chart — it is not an independent finding, and the chart does not state which version of the Artificial Analysis index its first row was read under.

And one omission is worth recording without interpreting it. The chart sets Grok 4.6 against Grok 4.5, GPT-5.6 Sol, and Claude Fable 5. Claude Opus 5 — which leads the Artificial Analysis Intelligence Index that the same chart cites, at 63.05 on the date we read it — is not in the comparison. We are reporting that it is absent, not why.

Four instruments give four different readings of Grok 4.6: the AA Intelligence Index v4.1.1 at 60.92 at high effort, LMArena at rank 46 with a pre-release flag, the Coding Agent Index v1.4 where it is not measured, and the vendor table where it leads 3 of 10 rows
Four instruments, four questions. A number from one of them does not transfer to a sentence about another.

Four Instruments, Four Different Questions

The short answer: the AA Intelligence Index, LMArena, the Coding Agent Index, and the vendor table measure four different things. A number from one of them cannot be carried into a sentence about another.

This is the reason a single headline about Grok 4.6 is always going to be wrong. Each instrument answers its own question:

  • Artificial Analysis Intelligence Index v4.1.1 asks: how does the model score across nine fixed evaluations run independently on dedicated hardware? Answer for Grok 4.6 at high: 60.92, level with GPT-5.6 Sol at max, behind Claude Opus 5 and Claude Fable 5.
  • LMArena asks: which response do blind human voters prefer? Answer: rank 46 with 1461.16 Elo on 3,473 votes, provisional, below Grok 4.5.
  • Coding Agent Index v1.4 asks: how does a harness-plus-model-plus-effort pair perform on three agentic coding benchmarks? Answer for Grok 4.6: not measured on August 22, 2026.
  • The vendor table asks: which numbers does SpaceXAI choose to publish, against competitor figures it describes as self-reported or public? Answer: Grok 4.6 leads 3 of 10 rows.

Put those four together and the defensible summary of Grok 4.6 on August 22, 2026 is this: a model that has reached the top group on one independent capability index at a cost per task that was materially lower on that date, though a re-reading six days later put that gap at 1.7 percent, that is unproven on the independent agentic-coding board, that human voters have so far preferred less than its own predecessor on thin data, and whose own launch chart it wins a minority of.

What the Tie Costs in Token Prices

The short answer: the cost-per-task figures above are one measurement; token list prices are another, and neither derives from the other. Both models bill in two tiers with different thresholds — Grok 4.6 crosses at 200,000 tokens, GPT-5.6 Sol at 272,000 — and OpenAI cut Sol’s rates on August 21, 2026 under a promotion whose stated duration differs between its own pages.

The other side of the tie moved too, and recently. On August 21, 2026 OpenAI cut the list price of GPT-5.6 Sol to $4.00 per million input tokens and $20.00 per million output tokens, down from the $5.00 and $30.00 it carried at its July launch. OpenAI’s API changelog describes that as “20% lower input pricing and 33% lower output pricing” — the percentages are the vendor’s own arithmetic, not ours — and states that “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.” It is a promotional rate, not a permanent repricing — though OpenAI describes how long it lasts differently on its announcement page than in its changelog, which is covered further down.

Those two numbers are not the whole rate card either, and here that matters more than usual. Sol bills in two tiers exactly as Grok 4.6 does — with a different threshold. OpenAI’s pricing page lists long-context Sol at $8.00 input and $30.00 output per million tokens, and its changelog defines long-context prompts as those exceeding 272,000 tokens. Grok 4.6 crosses into its own long tier at 200,000. Compare tier against tier, never a short-context rate against a long-context one.

ModelPrompt sizeInputOutput
Grok 4.6below 200,000 tokens$2.00 / 1M$6.00 / 1M
Grok 4.6200,000 tokens and above$4.00 / 1M$12.00 / 1M
GPT-5.6 Sol272,000 tokens and below$4.00 / 1M$20.00 / 1M
GPT-5.6 Solabove 272,000 tokens$8.00 / 1M$30.00 / 1M

Sources: SpaceXAI API pricing page and OpenAI API pricing page and changelog, read August 22, 2026. Sol’s rates are promotional at least through November 21, 2026.

Because the two thresholds differ, there is a band between 200,000 and 272,000 tokens where Grok 4.6 is already billing its long-context rate while Sol is still on its short one — and in that band the input rates meet exactly, at $4.00 for both, while Grok’s output stays at $12.00 against Sol’s $20.00. Below 200,000 tokens Grok lists at half Sol’s input rate and 30 percent of its output rate. Above 272,000 it is half the input and 40 percent of the output. A headline rate always hides a tier, and for the long-running agent workloads SpaceXAI is pitching this model at, the tier is the number that decides the bill.

One warning about mixing those numbers with the ones earlier in this article: token list prices and cost per task are two different measurements. The $0.9372 and $0.9530 figures cited earlier are what Artificial Analysis measured while running its index on real workloads, not a calculation from these rate cards. Neither number can be derived from the other.

Three Places the Primary Sources Disagree

The short answer: first-party sources give two different context-window figures for the same model, attach the same set of prices to two different concepts, and describe the duration of the same discount two different ways. We report each reading and do not arbitrate between them.

The context window: 500,000 or 256k

SpaceXAI's model documentation states a 500,000-token context window, and its Bedrock and Google Enterprise Agent Platform announcements repeat that figure. LMArena's leaderboard payload also records 500,000 for the entry. Cursor's documentation for the same model ID lists a context window of 256k.

Both are first-party sources — Cursor and SpaceXAI have been part of the same group since the Cursor acquisition closed. The most likely reading is that a harness can expose less context than the model supports, but the documentation does not say that, so we will not say it for them. What a developer should take away: 500,000 tokens is the figure SpaceXAI publishes for the API; 256k is the figure Cursor publishes for Grok 4.6 inside Cursor. Check the surface you are actually calling.

The second price tier: a speed variant or a context tier

Three first-party pages attach the identical numbers — $4.00 input, $1.00 cached, $12.00 output per million tokens — to three different explanations. The SpaceXAI pricing page presents them as the long-context tier, billed once a prompt reaches 200,000 tokens. Cursor's documentation presents them as the Fast variant, against standard on-demand usage at $2.00, $0.50, and $6.00. And the launch post says simply that "there is a fast variant which is twice the price."

There is no grok-4.6-fast line anywhere in the SpaceXAI price table, which lists grok-4.6, grok-build-0.1, grok-4.5, grok-4.3, and three grok-4.20 variants. So the same doubled rate card is described once as a speed tier and once as a context tier. If you are modeling costs, the practical consequence is that you have to know which surface you are billed through before you know what doubles your bill: a long prompt, or a fast lane.

How long the Sol discount lasts: a window, or a floor

The third disagreement is OpenAI’s, and it is about its own price cut. Two of its pages date the same discount differently, in the same week. The GPT-5.6 announcement page carries a dated note reading: “Update on August 21, 2026: OpenAI dropped the API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months.” The API pricing page and the API changelog both say instead: “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.”

Those are not the same statement. “For the next 3 months” describes a window that closes. “At least through” sets a floor the discount will not end before, and names no ceiling. One caps the promotion; the other guarantees a minimum and leaves the end open. Three months from August 21 does land on November 21, so the two are not numerically inconsistent — but they make opposite promises about what happens on that date.

We report both and do not arbitrate. The consequence for anyone budgeting is narrow and worth stating plainly: the real duration of the Sol discount is not determinable from OpenAI’s own pages. Treat November 21, 2026 as the earliest date on which the rate could change rather than the date it will, and re-read the pricing page after it rather than trusting any sentence written before — including this one.

Integration Notes Worth Knowing Before You Wire It In

The short answer: Grok 4.6 silently ignores logprobs rather than rejecting it, does not support the Batch API, and reprices an entire request at long-context rates once the prompt crosses 200,000 tokens.

These come from SpaceXAI's models documentation and the model page, and each has a failure mode attached:

  • logprobs and top_logprobs are not supported on grok-4.20 and newer, and the documentation is explicit that these fields "will be silently ignored if set." A rejected parameter surfaces in your error handling on the first call. A silently ignored one does not — any confidence-scoring or reranking logic downstream of those fields will simply receive nothing and keep running.
  • No Batch API. The model page lists Batch API as "Not supported." Offline bulk workloads built around batching on another provider do not port over as-is.
  • Long-context billing is all-or-nothing. Crossing 200,000 tokens reprices the whole request, not the excess. A retrieval step that occasionally over-fetches can double the cost of an otherwise ordinary call.
  • Effort is a cost lever, not a quality switch. With four levels available and high as the default, moving to xhigh spends more tokens for more deliberation — and, as covered above, what that buys is measured: about 31 percent more per index task, a lower Intelligence Index score, and a higher AA-Briefcase Elo.

Where You Can Actually Run It Today

The short answer: Grok 4.6 shipped in the SpaceXAI API and Cursor on August 12, reached GitHub Copilot on August 14, went generally available on Amazon Bedrock on August 19, and landed on Google Enterprise Agent Platform via Model Garden on August 21.

The rollout ran from launch to a third major enterprise platform in nine days. Each step is documented by the vendor:

Date (2026)SurfaceNote
August 12SpaceXAI API, Cursor, Grok BuildLaunch day; also listed on OpenRouter, Vercel, and Cloudflare
August 14GitHub CopilotModel picker across VS Code, Copilot CLI, and cloud agents
August 19Amazon BedrockGenerally available
August 21Google Enterprise Agent PlatformVia Model Garden

Sources: SpaceXAI announcement posts for the launch, GitHub Copilot, Amazon Bedrock, and Google Enterprise Agent Platform, all read August 22, 2026.

The Bedrock and Google posts both quote the same list rates as the API — $2.00 input, $0.50 cached input, $6.00 output per million tokens — so the pricing story is consistent across surfaces, with Cursor's Fast-variant framing as the exception noted above.

What This Actually Means

The competitive fact is that a SpaceXAI model has reached the same score as a frontier OpenAI model on a serious independent index. What that costs has moved fast: on August 22, 2026 Grok 4.6 ran the index for about a third less than GPT-5.6 Sol; at the August 28 reading the two are 1.7 percent apart, Sol having fallen and Grok 4.6 having risen. For teams whose workload looks like the nine evaluations in the Artificial Analysis Intelligence Index, the budget argument against Sol is now thin — against Claude Opus 5 at $2.337 per task, a 2.5x cost difference for 2.13 index points is a trade some teams will take and others will refuse, depending on how much the top of the range is worth to them.

The fact that sits next to it is that three other instruments do not confirm the story. The independent agentic-coding board has not run it. Human voters, on thin and provisional data, currently prefer its predecessor. And the company's own chart hands a majority of its rows to Claude Fable 5 while leaving out the model at the top of the index it cites.

None of that makes the tie fake. It makes it specific. Grok 4.6 is a model with one strong independent result, a measured ceiling at xhigh that scores below its own default, and, as of August 22, 2026, an unmeasured record on agentic coding. If the Coding Agent Index runs a Grok Build pairing with 4.6 and the LMArena entry matures past its pre-release flag, this page will be worth rereading. Until then, the tie on the index is the part of the story that is measured, independent, and still in SpaceXAI's favor.

Frequently Asked Questions

What is Grok 4.6?

Grok 4.6 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, released on August 12, 2026. It offers a 500,000-token context window, four configurable reasoning efforts (low, medium, high, and xhigh), and a knowledge cutoff of February 1, 2026. List pricing is $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for requests below 200,000 tokens.

Did Grok 4.6 really match GPT-5.6 Sol?

On the Artificial Analysis Intelligence Index v4.1.1, in a reading taken on August 22, 2026, Grok 4.6 at high effort scores 60.92 and GPT-5.6 Sol at max effort scores 60.93 — a difference of seven thousandths of a point. That is a tie on that index, at those configurations, on that date. The two models were not run at the same effort level, and the result does not transfer to other benchmarks. The decimals matter: SpaceXAI’s own chart rounds both models to 61, which hides the size of the gap, and rounding can equally make two models several tenths of a point apart print the same integer.

Why does the reasoning effort matter in this comparison?

Grok 4.6 is measured at high, which SpaceXAI ships as its default effort, while GPT-5.6 Sol and Claude Opus 5 are measured at max and Claude Fable 5 under the configuration Artificial Analysis labels with fallback. Artificial Analysis does publish scores for Grok 4.6 at xhigh, its top setting: 60.01 on the Intelligence Index, below the 60.92 the model posts at high, for about 31 percent more per task. The gap at maximum effort is therefore measured rather than unknown, and it runs the opposite way to the usual assumption. No source supports the claim that the two models were compared at equal settings.

How much cheaper is Grok 4.6 than GPT-5.6 Sol?

Barely, at the latest reading. Artificial Analysis measured $0.837 per Intelligence Index task for Grok 4.6 at high effort against $1.231 for GPT-5.6 Sol at max on August 22, 2026 — about 32 percent less. On a re-reading taken August 28, 2026 the same instrument gives $0.9372 for Grok 4.6 and $0.9530 for Sol, a difference of 1.7 percent. For reference, the same instrument measures $2.337 per task for Claude Opus 5 and $3.140 for Claude Fable 5, both unchanged between the two readings. Token list prices are a separate measurement and cannot be derived from cost per task, and both models bill in tiers: Grok 4.6 lists at $2.00 input and $6.00 output per million tokens below 200,000 tokens, and $4.00 and $12.00 at or above it, while GPT-5.6 Sol lists at $4.00 and $20.00 up to 272,000 tokens and $8.00 and $30.00 beyond. OpenAI cut Sol to those rates on August 21, 2026 and states on its pricing page and changelog that the promotional pricing is available at least through November 21, 2026, while its announcement page describes the same cut as applying for the next 3 months. The exact duration is not determinable from OpenAI's own pages.

Is Grok 4.6 the cheapest frontier model to run?

No. On the Artificial Analysis Cost per Task reading of August 28, 2026, GLM-5.3 at max effort runs the Intelligence Index for $0.683 per task, about 27 percent less than Grok 4.6 at $0.9372, while scoring 59.51 against 60.92. Kimi K3 at max runs it for $0.837, about 11 percent less. Grok 4.6 is the cheapest per task among the four models at the top of that index, not the cheapest overall.

Why does Grok 4.6 rank below Grok 4.5 on LMArena?

On LMArena's text leaderboard read on August 22, 2026, grok-4.6-high sits at rank 46 with an Elo of 1461.16 on 3,473 votes, while Grok 4.5 sits at rank 35 with 1470.29 on 22,030 votes. The newer entry is flagged pre_release, displays as preliminary, and has a rank interval spanning 25 to 69. Its confidence interval also overlaps Grok 4.5's, so the board does not cleanly separate the two models yet.

Is Grok 4.6 ranked on the Artificial Analysis Coding Agent Index?

No. On a reading of the Coding Agent Index v1.4 taken on August 22, 2026, none of the 55 harness-and-model pairs on the board involves Grok 4.6. The most recent SpaceXAI pairing is Grok Build with Grok 4.5 at high effort, which scores 0.641 at a mean cost of $2.44 per task. An absence is not a result: it means the independent agentic-coding board has not run the model yet.

How many benchmarks does Grok 4.6 win on SpaceXAI's own comparison table?

Three out of ten. On the table published in SpaceXAI's launch post, Grok 4.6 posts the leading figure on GDPVal-AA v2, AA-Briefcase, and Harvey LAB (Vals). Claude Fable 5 leads five rows and GPT-5.6 Sol leads two. SpaceXAI notes that third-party scores in that table are the best of self-reported or publicly available results, so these are vendor-published figures rather than independent measurements.

What is Grok 4.6's context window — 500,000 tokens or 256k?

Both figures come from first-party documentation. SpaceXAI publishes 500,000 tokens for grok-4.6 on its model page, and its Amazon Bedrock and Google Enterprise Agent Platform announcements repeat it. Cursor's documentation for the same model ID lists 256k inside its own harness. We report both rather than arbitrating between them: check the surface you are calling before sizing your prompts.

What is the Fast variant of Grok 4.6, and does it cost double?

The same rate card — $4.00 input, $1.00 cached input, and $12.00 output per million tokens — is described two different ways by first-party sources. SpaceXAI's pricing page presents it as the long-context tier, billed once a prompt reaches 200,000 tokens, while Cursor's documentation presents it as the Fast variant. There is no grok-4.6-fast entry in the SpaceXAI price table. The numbers are identical; the concepts attached to them are not.

Does Grok 4.6 support logprobs and the Batch API?

Neither. SpaceXAI's documentation states that logprobs and top_logprobs are not supported on grok-4.20 and newer, and that these fields will be silently ignored if set rather than rejected — so downstream confidence-scoring logic fails quietly. The grok-4.6 model page also lists Batch API support as "Not supported," which rules out offline bulk workloads built around batching.

Where can you use Grok 4.6 today?

Grok 4.6 shipped on August 12, 2026 in the SpaceXAI API, Cursor, and Grok Build, and was also listed on OpenRouter, Vercel, and Cloudflare. It reached GitHub Copilot on August 14, became generally available on Amazon Bedrock on August 19, and arrived on Google Enterprise Agent Platform through Model Garden on August 21. The model is served from the us-east-1 and us-west-2 regions with published limits of 150 requests per second and 50 million tokens per minute.

Sources

Published by Anthony Martinez, ThePlanetTools.ai. This article is editorial analysis of publicly documented information. ThePlanetTools.ai has no affiliation with SpaceXAI, OpenAI, Anthropic, or Cursor. Benchmark figures are attributed to their evaluators — Artificial Analysis and LMArena for independent measurements, SpaceXAI for vendor-published comparisons. Pricing and specifications come from official vendor documentation, read August 22, 2026.

Related Articles

Was this review helpful?
Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.