The Gist: SpaceXAI released Grok 4.6 on August 12, 2026, and on one independent instrument it now sits level with a top OpenAI model. On the Artificial Analysis Intelligence Index v4.1.1, in a reading taken on August 22, 2026, Grok 4.6 scores 60.92 against 60.93 for GPT-5.6 Sol — a gap of seven thousandths of a point — while costing $0.837 per index task against $1.231 for Sol, roughly 32 percent less. On that same index it still trails Claude Opus 5 at 63.05 and Claude Fable 5 at 62.07. The configuration matters and is not symmetrical: Grok 4.6 is measured at high, which SpaceXAI's own harness documents as its default effort, while GPT-5.6 Sol and Claude Opus 5 are measured at max and Claude Fable 5 under the configuration Artificial Analysis labels with fallback. Artificial Analysis publishes no score for Grok 4.6 at xhigh, so the gap at maximum effort is unknown rather than zero. Three other instruments read differently: on LMArena the model ranks 46th with 3,473 votes and a pre-release flag, below its own predecessor Grok 4.5 at 35th; on the Artificial Analysis Coding Agent Index v1.4 it is not measured at all; and on SpaceXAI's own comparison table it posts the leading figure in 3 of 10 rows.
Key Takeaways
- The tie is real, and it is one number on one index. Artificial Analysis Intelligence Index v4.1.1, read August 22, 2026: Grok 4.6 at
highscores 60.92, GPT-5.6 Sol atmaxscores 60.93. That is a tie in any practical sense — and it is a composite of nine evaluations, not a verdict on everything a model does. - The price argument is stronger than the score argument. Artificial Analysis measures $0.837 per index task for Grok 4.6 against $1.231 for GPT-5.6 Sol, $2.337 for Claude Opus 5, and $3.140 for Claude Fable 5. Grok 4.6 delivers the same index score as Sol for about a third less. It is not, however, the cheapest thing on the board: GLM-5.3 at
maxruns the same index for $0.683. - The effort levels are not matched, and nobody should pretend otherwise. SpaceXAI documents four reasoning efforts for Grok 4.6 — low, medium, high, and xhigh — and Cursor's documentation names
highas the default. Artificial Analysis publishes noxhighfigure. Anyone writing "at equal settings" about this comparison is inventing a measurement that does not exist. - On human preference it ranks below the model it replaces. LMArena's text leaderboard, read August 22, 2026, places
grok-4.6-highat rank 46 with an Elo of 1461.16, against rank 35 and 1470.29 for Grok 4.5. The reserve is large: 3,473 votes, apre_releaseflag, and a rank interval spanning 25 to 69. - Its own launch chart leaves out the model currently leading the index it cites. SpaceXAI's comparison table sets Grok 4.6 against Grok 4.5, GPT-5.6 Sol, and Claude Fable 5. Claude Opus 5, which tops the Artificial Analysis Intelligence Index at 63.05 on the same date, does not appear in it.
What SpaceXAI Actually Shipped on August 12
SpaceXAI announced Grok 4.6 on August 12, 2026, positioning it as a successor to Grok 4.5 aimed at long-running agents and, in the company's words, "more ambitious interactive and visual work." The company is the entity formerly known as xAI, which completed its rebrand earlier in 2026; we covered that change in xAI Is Officially SpaceXAI Now, and the model line kept the Grok name.
The hard specifications come from the vendor's model page. According to SpaceXAI's documentation for grok-4.6, the model exposes a 500,000-token context window, accepts text and image input and returns text, and supports function calling, structured outputs, and configurable reasoning. Its knowledge cutoff is February 1, 2026. It is served from two regions, us-east-1 and us-west-2, with published rate limits of 150 requests per second and 50 million tokens per minute.
List pricing, from SpaceXAI's API pricing page, has two tiers. Short context bills $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens. Long context — which the page defines as any request whose prompt reaches 200,000 tokens — bills $4.00, $1.00, and $12.00 respectively.
| grok-4.6 list price | Input | Cached input | Output |
|---|---|---|---|
| Short context (below 200k tokens) | $2.00 / 1M | $0.50 / 1M | $6.00 / 1M |
| Long context (prompt ≥ 200k tokens) | $4.00 / 1M | $1.00 / 1M | $12.00 / 1M |
Source: SpaceXAI API pricing page, read August 22, 2026. The long-context threshold is the instrument here — see the section on where the sources disagree.
One detail in that table is easy to miss and expensive to learn late: the pricing page states that models with long-context pricing bill the long-context rates for all tokens in a request once its prompt reaches the threshold. Crossing 200,000 tokens does not surcharge the overflow. It reprices the entire call.
Worth noting against its predecessor: input and output rates are identical to Grok 4.5 at $2.00 and $6.00, but the cached-input rate moved from $0.30 to $0.50 per million tokens. For a cache-heavy agent loop, the newer model is the more expensive one on that line.
The Index Where Grok 4.6 Ties a Frontier OpenAI Model
The short answer: on the Artificial Analysis Intelligence Index v4.1.1, read on August 22, 2026, Grok 4.6 at high effort scores 60.92 and GPT-5.6 Sol at max scores 60.93. The two are level to within seven thousandths of a point, and Grok 4.6 runs the same battery for about a third less money.
The index is a composite. Artificial Analysis defines v4.1.1 as nine evaluations combined: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. The evaluator states that results are measured independently on dedicated hardware. Here is the top of that board on the day we read it.
That version number is not decoration, because the instrument itself moves. Artificial Analysis records what changed in v4.1.1: 𝜏³-Banking was updated to v1.0.1, and the graders for Humanity’s Last Exam, AA-LCR and AA-Omniscience were upgraded to GPT-5.6 Luna (medium). The practical consequence is worth stating plainly: a score read under an earlier version is not comparable to one read under v4.1.1. When a model’s number moves between two readings taken months apart, the model may not have changed at all — the measurement did. Every Artificial Analysis figure on this page comes from one reading of v4.1.1, taken on August 22, 2026, which is why no score here is presented as a gain or a loss over time.
| Model (effort as measured) | AA Intelligence Index v4.1.1 | Cost per index task |
|---|---|---|
| Claude Opus 5 (max) | 63.05 | $2.337 |
| Claude Fable 5 (with fallback) | 62.07 | $3.140 |
| GPT-5.6 Sol (max) | 60.93 | $1.231 |
| Grok 4.6 (high) | 60.92 | $0.837 |
| Kimi K3 (max) | 59.70 | $0.837 |
| GLM-5.3 (max) | 59.51 | $0.683 |
Source: Artificial Analysis, Intelligence Index v4.1.1 and Cost per Task datasets, read August 22, 2026. Effort levels are as the evaluator ran them and are not uniform across rows.
Two things in that table deserve to be said out loud rather than left for the reader to infer. First, the models above Grok 4.6 are still above it: Claude Opus 5 leads by 2.13 points and Claude Fable 5 by 1.15 points on this index. Second, the cost column does not make Grok 4.6 the budget option in any absolute sense. Kimi K3 runs the same battery for the same $0.837, and GLM-5.3 runs it for $0.683 — around 18 percent less than Grok 4.6 — while scoring 1.4 points lower. The honest framing is narrower and still favorable: among the four models at the top of this index, Grok 4.6 is the cheapest per task by a wide margin.
Artificial Analysis also measures a median output speed of 68.7 tokens per second for Grok 4.6 at high, which places it between GPT-5.6 Sol at 69.98 and Claude Opus 5 at 58.39 on the same reading. This is not a fast model relative to the wider field — Gemini 3.7 Flash at high measures 389.5 on the same instrument.
The Asterisk That Decides How Much the Tie Is Worth
The short answer: the two tied entries were not run at the same reasoning effort. Grok 4.6 is measured at high, its default; GPT-5.6 Sol is measured at max. Artificial Analysis publishes no Grok 4.6 score at xhigh, so what the model would do at its own ceiling is unmeasured, not equal.
SpaceXAI documents four reasoning efforts for Grok 4.6 — low, medium, high, and xhigh — in its Amazon Bedrock announcement, and Cursor's model documentation states plainly that high is the default. So the index is comparing a model at its shipped default against competitors measured under settings other than its own — max for GPT-5.6 Sol and Claude Opus 5, and for Claude Fable 5 the configuration Artificial Analysis labels with fallback.
That cuts in an interesting direction, and it is worth being precise about which. It is not evidence that Grok 4.6 is quietly better than the tie suggests — nobody has published that measurement. It is evidence that the comparison is asymmetric in a way most coverage will flatten. Higher effort generally costs more tokens and more money, so a high-effort score at $0.837 per task and a max-effort score at $1.231 per task are not the same product decision even when the two numbers land on top of each other.
The only defensible statement is the narrow one: at the configurations Artificial Analysis actually ran on August 22, 2026, the two models score the same on this index, and Grok 4.6 does it for less. Anything stronger — "matches at equal settings," "beats it when you push it" — is a claim about a measurement that does not exist.
The Board Where It Ranks Below Its Own Predecessor
The short answer: on LMArena's text leaderboard, read August 22, 2026, grok-4.6-high sits at rank 46 with an Elo of 1461.16, while Grok 4.5 sits at rank 35 with 1470.29. The newer model is placed below the older one — on a provisional entry with 3,473 votes and a pre-release flag.
LMArena measures something different from Artificial Analysis: blind human preference between two anonymous model responses, aggregated into an Elo rating. It is not a capability battery, and a score here does not transfer to the index above or the other way around.
| Entry | Rank | Elo | Confidence interval | Votes |
|---|---|---|---|---|
| claude-fable-5 | 1 | 1507.77 | 1502.58 – 1512.95 | 24,331 |
| claude-opus-5-high | 7 | 1493 | ±5 | 27,610 |
| gpt-5.6-sol-xhigh | 16 | 1482.26 | 1476.86 – 1487.66 | 19,541 |
| grok-4.5 | 35 | 1470.29 | 1465.13 – 1475.45 | 22,030 |
| grok-4.6-high | 46 | 1461.16 | 1451.07 – 1471.26 | 3,473 |
Source: LMArena text leaderboard payload, read August 22, 2026. The grok-4.6-high entry carries a pre_release flag and displays as preliminary.
The reserve on that row is not a formality, and we would rather state it than bury it. LMArena marks the entry pre_release and displays it as preliminary. Its rank interval runs from 25 to 69 — a 44-place span. And its confidence interval, 1451.07 to 1471.26, overlaps Grok 4.5's interval of 1465.13 to 1475.45 across the range 1465.13 to 1471.26. In plain terms: the board currently places Grok 4.6 below Grok 4.5, but with this many votes it does not cleanly separate the two. That ordering can move.
What is harder to wave away is the company it keeps. Five SpaceXAI entries rank above grok-4.6-high on the same board — grok-4.20-beta1 at 25, grok-4.20-beta-0309-reasoning at 33, grok-4.5 at 35, grok-4.20-multi-agent-beta-0309 at 36, and grok-4.1-thinking at 42. For a flagship release, being placed behind five of your own earlier entries on a human-preference board is a signal about conversational preference, even a provisional one — and it is not something the launch chart shows.
The Board Where It Is Not Measured at All
The short answer: on the Artificial Analysis Coding Agent Index v1.4, read August 22, 2026, there is no Grok 4.6 entry. Across the 55 harness-and-model pairs the board carries, the most recent SpaceXAI pairing is "Grok Build - Grok 4.5 (high)."
This matters because SpaceXAI presents Grok 4.6 as a model for long-running agents and coding, and the Coding Agent Index is the instrument built for exactly that. Version 1.4 combines three benchmarks — DeepSWE, SWE-Atlas-QnA, and Terminal-Bench v2.1 — and, crucially, it scores pairs: a harness plus a model plus an effort level, not a model in the abstract. "Claude Code - Opus 5 (xhigh)" and "Codex - GPT-5.6 Sol (max)" are separate entries because the harness is part of what is being measured.
One caveat on our own method, since it applies to anyone checking this: the structured data published in that page's markup exposes only nine display rows. Concluding "not ranked" from those nine would be wrong. We read the full board payload, counted 55 pairs, and found no Grok 4.6 pairing at any effort. Grok Build with Grok 4.5 at high scores 0.641 there at a mean $2.44 per task, against 0.681 for Claude Code with Opus 5 at xhigh at $8.17 per task. We covered the harness itself in SpaceXAI Open-Sources Grok Build.
An absence is not a result. It does not mean Grok 4.6 would score badly on that index; it means that on August 22, 2026 the independent agentic-coding board has not yet run it, so anyone citing Grok 4.6's agentic-coding standing is citing the vendor, not an evaluator.
What the Vendor's Own Table Says — and What It Leaves Out
The short answer: on SpaceXAI's own comparison table, Grok 4.6 posts the leading figure in 3 of 10 rows. These are vendor-published numbers, and the company states that competitor figures are self-reported or drawn from public results.
The table below is reproduced from SpaceXAI's launch post, which carries the note: "Third-party model scores are the best of self-reported or publicly available results." That footnote is the instrument. These are not independent measurements, and we present them as the company's claims.
| Evaluation (vendor-published) | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
Source: SpaceXAI launch post, read August 22, 2026. Bold marks the leading figure as the vendor marks it. Competitor figures are self-reported or publicly available results per the vendor's own note.
Three observations, all factual. Grok 4.6 leads 3 of the 10 rows — GDPVal-AA v2, AA-Briefcase, and Harvey LAB. Claude Fable 5 leads five of them and GPT-5.6 Sol leads two, on a chart the vendor chose. The gap is widest where the task is agentic and terminal-shaped: 26 percent against 34.6 percent on Terminal-Bench v3.0, and 65.9 against 73 on DeepSWE v1.1, both against GPT-5.6 Sol.
One reconciliation, because the two tables on this page disagree at a glance. The vendor chart rounds to whole numbers. It prints 61 for both Grok 4.6 and GPT-5.6 Sol and 62 for Claude Fable 5, where the Artificial Analysis figures for that same index are 60.92, 60.93 and 62.07. Rounding is why the vendor’s chart can show a flat tie at 61. It cuts the other way too: on a rounded board, two models separated by several tenths of a point can print the same integer and look level when they are not. That is the reason every index figure in this article is given to two decimals.
The improvement over its own predecessor is the cleanest story in the table: as the vendor presents it, Grok 4.6 is ahead of Grok 4.5 in all ten rows, by 227 points on GDPVal-AA v2 and by more than 10 points on Terminal-Bench v3.0 and DeepSWE v1.1. That comparison is the vendor’s own, drawn on the vendor’s chart — it is not an independent finding, and the chart does not state which version of the Artificial Analysis index its first row was read under.
And one omission is worth recording without interpreting it. The chart sets Grok 4.6 against Grok 4.5, GPT-5.6 Sol, and Claude Fable 5. Claude Opus 5 — which leads the Artificial Analysis Intelligence Index that the same chart cites, at 63.05 on the date we read it — is not in the comparison. We are reporting that it is absent, not why.
Four Instruments, Four Different Questions
The short answer: the AA Intelligence Index, LMArena, the Coding Agent Index, and the vendor table measure four different things. A number from one of them cannot be carried into a sentence about another.
This is the reason a single headline about Grok 4.6 is always going to be wrong. Each instrument answers its own question:
- Artificial Analysis Intelligence Index v4.1.1 asks: how does the model score across nine fixed evaluations run independently on dedicated hardware? Answer for Grok 4.6 at
high: 60.92, level with GPT-5.6 Sol atmax, behind Claude Opus 5 and Claude Fable 5. - LMArena asks: which response do blind human voters prefer? Answer: rank 46 with 1461.16 Elo on 3,473 votes, provisional, below Grok 4.5.
- Coding Agent Index v1.4 asks: how does a harness-plus-model-plus-effort pair perform on three agentic coding benchmarks? Answer for Grok 4.6: not measured on August 22, 2026.
- The vendor table asks: which numbers does SpaceXAI choose to publish, against competitor figures it describes as self-reported or public? Answer: Grok 4.6 leads 3 of 10 rows.
Put those four together and the defensible summary of Grok 4.6 on August 22, 2026 is this: a model that has reached the top group on one independent capability index at a materially lower cost per task, that is unproven on the independent agentic-coding board, that human voters have so far preferred less than its own predecessor on thin data, and whose own launch chart it wins a minority of.
What the Tie Costs in Token Prices
The short answer: the cost-per-task figures above are one measurement; token list prices are another, and neither derives from the other. Both models bill in two tiers with different thresholds — Grok 4.6 crosses at 200,000 tokens, GPT-5.6 Sol at 272,000 — and OpenAI cut Sol’s rates on August 21, 2026 under a promotion whose stated duration differs between its own pages.
The other side of the tie moved too, and recently. On August 21, 2026 OpenAI cut the list price of GPT-5.6 Sol to $4.00 per million input tokens and $20.00 per million output tokens, down from the $5.00 and $30.00 it carried at its July launch. OpenAI’s API changelog describes that as “20% lower input pricing and 33% lower output pricing” — the percentages are the vendor’s own arithmetic, not ours — and states that “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.” It is a promotional rate, not a permanent repricing — though OpenAI describes how long it lasts differently on its announcement page than in its changelog, which is covered further down.
Those two numbers are not the whole rate card either, and here that matters more than usual. Sol bills in two tiers exactly as Grok 4.6 does — with a different threshold. OpenAI’s pricing page lists long-context Sol at $8.00 input and $30.00 output per million tokens, and its changelog defines long-context prompts as those exceeding 272,000 tokens. Grok 4.6 crosses into its own long tier at 200,000. Compare tier against tier, never a short-context rate against a long-context one.
| Model | Prompt size | Input | Output |
|---|---|---|---|
| Grok 4.6 | below 200,000 tokens | $2.00 / 1M | $6.00 / 1M |
| Grok 4.6 | 200,000 tokens and above | $4.00 / 1M | $12.00 / 1M |
| GPT-5.6 Sol | 272,000 tokens and below | $4.00 / 1M | $20.00 / 1M |
| GPT-5.6 Sol | above 272,000 tokens | $8.00 / 1M | $30.00 / 1M |
Sources: SpaceXAI API pricing page and OpenAI API pricing page and changelog, read August 22, 2026. Sol’s rates are promotional at least through November 21, 2026.
Because the two thresholds differ, there is a band between 200,000 and 272,000 tokens where Grok 4.6 is already billing its long-context rate while Sol is still on its short one — and in that band the input rates meet exactly, at $4.00 for both, while Grok’s output stays at $12.00 against Sol’s $20.00. Below 200,000 tokens Grok lists at half Sol’s input rate and 30 percent of its output rate. Above 272,000 it is half the input and 40 percent of the output. A headline rate always hides a tier, and for the long-running agent workloads SpaceXAI is pitching this model at, the tier is the number that decides the bill.
One warning about mixing those numbers with the ones earlier in this article: token list prices and cost per task are two different measurements. The $0.837 and $1.231 figures cited earlier are what Artificial Analysis measured while running its index on real workloads, not a calculation from these rate cards. Neither number can be derived from the other.
Three Places the Primary Sources Disagree
The short answer: first-party sources give two different context-window figures for the same model, attach the same set of prices to two different concepts, and describe the duration of the same discount two different ways. We report each reading and do not arbitrate between them.
The context window: 500,000 or 256k
SpaceXAI's model documentation states a 500,000-token context window, and its Bedrock and Google Enterprise Agent Platform announcements repeat that figure. LMArena's leaderboard payload also records 500,000 for the entry. Cursor's documentation for the same model ID lists a context window of 256k.
Both are first-party sources — Cursor and SpaceXAI have been part of the same group since the Cursor acquisition closed. The most likely reading is that a harness can expose less context than the model supports, but the documentation does not say that, so we will not say it for them. What a developer should take away: 500,000 tokens is the figure SpaceXAI publishes for the API; 256k is the figure Cursor publishes for Grok 4.6 inside Cursor. Check the surface you are actually calling.
The second price tier: a speed variant or a context tier
Three first-party pages attach the identical numbers — $4.00 input, $1.00 cached, $12.00 output per million tokens — to three different explanations. The SpaceXAI pricing page presents them as the long-context tier, billed once a prompt reaches 200,000 tokens. Cursor's documentation presents them as the Fast variant, against standard on-demand usage at $2.00, $0.50, and $6.00. And the launch post says simply that "there is a fast variant which is twice the price."
There is no grok-4.6-fast line anywhere in the SpaceXAI price table, which lists grok-4.6, grok-build-0.1, grok-4.5, grok-4.3, and three grok-4.20 variants. So the same doubled rate card is described once as a speed tier and once as a context tier. If you are modeling costs, the practical consequence is that you have to know which surface you are billed through before you know what doubles your bill: a long prompt, or a fast lane.
How long the Sol discount lasts: a window, or a floor
The third disagreement is OpenAI’s, and it is about its own price cut. Two of its pages date the same discount differently, in the same week. The GPT-5.6 announcement page carries a dated note reading: “Update on August 21, 2026: OpenAI dropped the API and credit pricing of GPT-5.6 Sol by over 20% for the next 3 months.” The API pricing page and the API changelog both say instead: “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.”
Those are not the same statement. “For the next 3 months” describes a window that closes. “At least through” sets a floor the discount will not end before, and names no ceiling. One caps the promotion; the other guarantees a minimum and leaves the end open. Three months from August 21 does land on November 21, so the two are not numerically inconsistent — but they make opposite promises about what happens on that date.
We report both and do not arbitrate. The consequence for anyone budgeting is narrow and worth stating plainly: the real duration of the Sol discount is not determinable from OpenAI’s own pages. Treat November 21, 2026 as the earliest date on which the rate could change rather than the date it will, and re-read the pricing page after it rather than trusting any sentence written before — including this one.
Integration Notes Worth Knowing Before You Wire It In
The short answer: Grok 4.6 silently ignores logprobs rather than rejecting it, does not support the Batch API, and reprices an entire request at long-context rates once the prompt crosses 200,000 tokens.
These come from SpaceXAI's models documentation and the model page, and each has a failure mode attached:
logprobsandtop_logprobsare not supported on grok-4.20 and newer, and the documentation is explicit that these fields "will be silently ignored if set." A rejected parameter surfaces in your error handling on the first call. A silently ignored one does not — any confidence-scoring or reranking logic downstream of those fields will simply receive nothing and keep running.- No Batch API. The model page lists Batch API as "Not supported." Offline bulk workloads built around batching on another provider do not port over as-is.
- Long-context billing is all-or-nothing. Crossing 200,000 tokens reprices the whole request, not the excess. A retrieval step that occasionally over-fetches can double the cost of an otherwise ordinary call.
- Effort is a cost lever, not a quality switch. With four levels available and
highas the default, moving toxhighspends more tokens for more deliberation — and, as covered above, no independent index has published what that buys on this model.
Where You Can Actually Run It Today
The short answer: Grok 4.6 shipped in the SpaceXAI API and Cursor on August 12, reached GitHub Copilot on August 14, went generally available on Amazon Bedrock on August 19, and landed on Google Enterprise Agent Platform via Model Garden on August 21.
The rollout ran from launch to a third major enterprise platform in nine days. Each step is documented by the vendor:
| Date (2026) | Surface | Note |
|---|---|---|
| August 12 | SpaceXAI API, Cursor, Grok Build | Launch day; also listed on OpenRouter, Vercel, and Cloudflare |
| August 14 | GitHub Copilot | Model picker across VS Code, Copilot CLI, and cloud agents |
| August 19 | Amazon Bedrock | Generally available |
| August 21 | Google Enterprise Agent Platform | Via Model Garden |
Sources: SpaceXAI announcement posts for the launch, GitHub Copilot, Amazon Bedrock, and Google Enterprise Agent Platform, all read August 22, 2026.
The Bedrock and Google posts both quote the same list rates as the API — $2.00 input, $0.50 cached input, $6.00 output per million tokens — so the pricing story is consistent across surfaces, with Cursor's Fast-variant framing as the exception noted above.
What This Actually Means
The competitive fact is that a SpaceXAI model has reached the same score as a frontier OpenAI model on a serious independent index, and does it for about a third less per task. For teams whose workload looks like the nine evaluations in the Artificial Analysis Intelligence Index, that is a real budget argument — and against Claude Opus 5 at $2.337 per task, a 2.8x cost difference for 2.13 index points is a trade some teams will take and others will refuse, depending on how much the top of the range is worth to them.
The fact that sits next to it is that three other instruments do not confirm the story. The independent agentic-coding board has not run it. Human voters, on thin and provisional data, currently prefer its predecessor. And the company's own chart hands a majority of its rows to Claude Fable 5 while leaving out the model at the top of the index it cites.
None of that makes the tie fake. It makes it specific. Grok 4.6 is, on August 22, 2026, a model with one strong independent result at an aggressive price, an unmeasured ceiling at xhigh, and an unmeasured record on agentic coding — with three of those four statements being about what has not been tested yet. If the Coding Agent Index runs a Grok Build pairing with 4.6 and the LMArena entry matures past its pre-release flag, this page will be worth rereading. Until then, the cost-per-task line is the part of the story that is measured, independent, and in SpaceXAI's favor.
Frequently Asked Questions
What is Grok 4.6?
Grok 4.6 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, released on August 12, 2026. It offers a 500,000-token context window, four configurable reasoning efforts (low, medium, high, and xhigh), and a knowledge cutoff of February 1, 2026. List pricing is $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for requests below 200,000 tokens.
Did Grok 4.6 really match GPT-5.6 Sol?
On the Artificial Analysis Intelligence Index v4.1.1, in a reading taken on August 22, 2026, Grok 4.6 at high effort scores 60.92 and GPT-5.6 Sol at max effort scores 60.93 — a difference of seven thousandths of a point. That is a tie on that index, at those configurations, on that date. The two models were not run at the same effort level, and the result does not transfer to other benchmarks. The decimals matter: SpaceXAI’s own chart rounds both models to 61, which hides the size of the gap, and rounding can equally make two models several tenths of a point apart print the same integer.
Why does the reasoning effort matter in this comparison?
Grok 4.6 is measured at high, which SpaceXAI ships as its default effort, while GPT-5.6 Sol and Claude Opus 5 are measured at max and Claude Fable 5 under the configuration Artificial Analysis labels with fallback. Artificial Analysis publishes no score for Grok 4.6 at xhigh, its top setting. That means the gap at maximum effort is unknown rather than zero, and no source supports the claim that the models were compared at equal settings.
How much cheaper is Grok 4.6 than GPT-5.6 Sol?
Artificial Analysis measures a cost of $0.837 per Intelligence Index task for Grok 4.6 at high effort, against $1.231 for GPT-5.6 Sol at max effort, in a reading taken on August 22, 2026. That is about 32 percent less. For reference, the same instrument measures $2.337 per task for Claude Opus 5 and $3.140 for Claude Fable 5. Token list prices are a separate measurement and cannot be derived from cost per task, and both models bill in tiers: Grok 4.6 lists at $2.00 input and $6.00 output per million tokens below 200,000 tokens, and $4.00 and $12.00 at or above it, while GPT-5.6 Sol lists at $4.00 and $20.00 up to 272,000 tokens and $8.00 and $30.00 beyond. OpenAI cut Sol to those rates on August 21, 2026 and states on its pricing page and changelog that the promotional pricing is available at least through November 21, 2026, while its announcement page describes the same cut as applying for the next 3 months. The exact duration is not determinable from OpenAI's own pages.
Is Grok 4.6 the cheapest frontier model to run?
No. On the Artificial Analysis Cost per Task reading of August 22, 2026, GLM-5.3 at max effort runs the Intelligence Index for $0.683 per task, about 18 percent less than Grok 4.6, while scoring 59.51 against 60.92. Kimi K3 at max costs the same $0.837. Grok 4.6 is the cheapest per task among the four models at the top of that index, not the cheapest overall.
Why does Grok 4.6 rank below Grok 4.5 on LMArena?
On LMArena's text leaderboard read on August 22, 2026, grok-4.6-high sits at rank 46 with an Elo of 1461.16 on 3,473 votes, while Grok 4.5 sits at rank 35 with 1470.29 on 22,030 votes. The newer entry is flagged pre_release, displays as preliminary, and has a rank interval spanning 25 to 69. Its confidence interval also overlaps Grok 4.5's, so the board does not cleanly separate the two models yet.
Is Grok 4.6 ranked on the Artificial Analysis Coding Agent Index?
No. On a reading of the Coding Agent Index v1.4 taken on August 22, 2026, none of the 55 harness-and-model pairs on the board involves Grok 4.6. The most recent SpaceXAI pairing is Grok Build with Grok 4.5 at high effort, which scores 0.641 at a mean cost of $2.44 per task. An absence is not a result: it means the independent agentic-coding board has not run the model yet.
How many benchmarks does Grok 4.6 win on SpaceXAI's own comparison table?
Three out of ten. On the table published in SpaceXAI's launch post, Grok 4.6 posts the leading figure on GDPVal-AA v2, AA-Briefcase, and Harvey LAB (Vals). Claude Fable 5 leads five rows and GPT-5.6 Sol leads two. SpaceXAI notes that third-party scores in that table are the best of self-reported or publicly available results, so these are vendor-published figures rather than independent measurements.
What is Grok 4.6's context window — 500,000 tokens or 256k?
Both figures come from first-party documentation. SpaceXAI publishes 500,000 tokens for grok-4.6 on its model page, and its Amazon Bedrock and Google Enterprise Agent Platform announcements repeat it. Cursor's documentation for the same model ID lists 256k inside its own harness. We report both rather than arbitrating between them: check the surface you are calling before sizing your prompts.
What is the Fast variant of Grok 4.6, and does it cost double?
The same rate card — $4.00 input, $1.00 cached input, and $12.00 output per million tokens — is described two different ways by first-party sources. SpaceXAI's pricing page presents it as the long-context tier, billed once a prompt reaches 200,000 tokens, while Cursor's documentation presents it as the Fast variant. There is no grok-4.6-fast entry in the SpaceXAI price table. The numbers are identical; the concepts attached to them are not.
Does Grok 4.6 support logprobs and the Batch API?
Neither. SpaceXAI's documentation states that logprobs and top_logprobs are not supported on grok-4.20 and newer, and that these fields will be silently ignored if set rather than rejected — so downstream confidence-scoring logic fails quietly. The grok-4.6 model page also lists Batch API support as "Not supported," which rules out offline bulk workloads built around batching.
Where can you use Grok 4.6 today?
Grok 4.6 shipped on August 12, 2026 in the SpaceXAI API, Cursor, and Grok Build, and was also listed on OpenRouter, Vercel, and Cloudflare. It reached GitHub Copilot on August 14, became generally available on Amazon Bedrock on August 19, and arrived on Google Enterprise Agent Platform through Model Garden on August 21. The model is served from the us-east-1 and us-west-2 regions with published limits of 150 requests per second and 50 million tokens per minute.
Sources
- SpaceXAI — Introducing Grok 4.6 (August 12, 2026)
- SpaceXAI — grok-4.6 model page (specifications, regions, rate limits, Batch API)
- SpaceXAI — Models documentation (knowledge cutoff, logprobs behavior)
- SpaceXAI — API pricing (short and long context rates)
- Artificial Analysis — Model leaderboards (Intelligence Index v4.1.1, Cost per Task, Output Speed; read August 22, 2026)
- Artificial Analysis — Coding Agent Index (v1.4; read August 22, 2026)
- LMArena — Text leaderboard (read August 22, 2026)
- Cursor — Grok 4.6 model documentation (effort levels, context window, Fast variant)
- OpenAI — GPT-5.6 announcement (dated update of August 21, 2026 on the Sol price cut)
- OpenAI — Models and pricing documentation (current GPT-5.6 list rates; read August 22, 2026)
- OpenAI — API pricing (GPT-5.6 short and long context tiers; read August 22, 2026)
- OpenAI — API changelog (August 21, 2026 entry on the Sol price cut; August 5, 2026 entry defining long-context prompts as exceeding 272K tokens)
- SpaceXAI — Grok 4.6 in GitHub Copilot (August 14, 2026)
- SpaceXAI — Grok 4.6 on Amazon Bedrock (August 19, 2026)
- SpaceXAI — Grok 4.6 on Google Enterprise Agent Platform (August 21, 2026)
Published by Anthony Martinez, ThePlanetTools.ai. This article is editorial analysis of publicly documented information. ThePlanetTools.ai has no affiliation with SpaceXAI, OpenAI, Anthropic, or Cursor. Benchmark figures are attributed to their evaluators — Artificial Analysis and LMArena for independent measurements, SpaceXAI for vendor-published comparisons. Pricing and specifications come from official vendor documentation, read August 22, 2026.



