Skip to content
G

Grok 4.6

SpaceXAI's frontier model for long-running agents — 500K context, four reasoning efforts, and the lowest list input rate at the top of the index.

8.0/10
Last updated August 29, 2026
Author
Anthony M.
49 min readVerified August 29, 2026Tested hands-on

Quick Summary

Grok 4.6 is SpaceXAI’s frontier model for coding and agentic work, released August 12, 2026. Score 8.0/10. From $2.00 per million input tokens, 500K context. Its default high effort outscores its top xhigh setting on the AA Intelligence Index.

Grok 4.6 tool review — SpaceXAI’s frontier model scored 8.0 out of 10, with a 500K context window and four reasoning efforts
Grok 4.6 — SpaceXAI's frontier model for long-running agents. Our score: 8.0 out of 10.

Grok 4.6 is SpaceXAI's frontier model for coding, agentic work, and knowledge tasks, released on August 12, 2026. We score it 8.0 out of 10. It ships a 500,000-token context window, takes text and image input and returns text only, and exposes four reasoning efforts — low, medium, high, and xhigh — with reasoning that cannot be switched off. List pricing starts at $2.00 per million input tokens and $6.00 per million output tokens, doubling once a prompt reaches 200,000 tokens. Independent measurement shows its default high setting scoring higher than its top xhigh setting on one capability index.

Verdict First: 8.0 out of 10

Grok 4.6 is a strong buy for teams that need frontier-class reasoning on a tight per-task budget and can live inside a US-shaped deployment. It is a poor fit for teams with EU data-residency requirements, offline batch pipelines, or any dependency on token-level probabilities.

ScoreRatingWhy
Overall8.0 / 10Frontier-tier capability at the lowest list price in its class, held back by documentation that contradicts itself on four separate points
Features8.5 / 10500,000-token context, four reasoning efforts, function calling, structured outputs, image input — but no Batch API, no logprobs, no image output
Ease of use7.5 / 10OpenAI-shaped Responses and Chat Completions APIs, but the default effort differs by surface and prompt caching silently degrades without an explicit routing key
Value9.0 / 10$2.00 per million input tokens against $4.00 for GPT-5.6 Sol and $5.00 for Claude Opus 5, at a measured index score within a hundredth of a point of Sol
Support7.0 / 10Dense, well-organized developer docs — that state two different knowledge cutoffs and two different default efforts for the same model

Scores are ours. Every fact below is attributed to a primary source we opened on August 28, 2026.

Pricing at a Glance

TierInputCached inputOutput
Prompt below 200,000 tokens$2.00 / 1M$0.50 / 1M$6.00 / 1M
Prompt at or above 200,000 tokens$4.00 / 1M$1.00 / 1M$12.00 / 1M
Priority Processing2x the standard rate across all token types
Blocked request$0.05 usage guideline violation fee per request
Batch APINot supported

Source: SpaceXAI API pricing, read August 28, 2026.

The single most expensive thing to learn late: the pricing page states that "requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request." Crossing 200,000 tokens does not surcharge the overflow. It reprices the entire call. A retrieval step that occasionally over-fetches by a thousand tokens does not add a thousand tokens' worth of cost — it doubles the bill for that request.

For context on where that sits in the market: GPT-5.6 Sol lists at $4.00 per million input tokens and Claude Opus 5 at $5.00, both against Grok 4.6's $2.00 — with the caveat that OpenAI describes Sol's current rate as promotional pricing available at least through November 21, 2026. We covered the index result that put those three models within striking distance of each other in Grok 4.6 Ties GPT-5.6 Sol on the AA Intelligence Index.

How We Assessed Grok 4.6

We have not run Grok 4.6 through a controlled hands-on evaluation, and we are not going to pretend otherwise. This review is research-led: every specification, price, limit, and score below was read directly from a primary source — SpaceXAI's developer documentation and announcement posts, Amazon's Bedrock model card, Cursor's model documentation, and the raw leaderboard payloads published by Artificial Analysis and Arena — on August 28, 2026. Where two primary sources disagree, we print both readings and name the surface each one describes, rather than picking the one that reads better. Where a figure could not be verified, we say so instead of filling the gap.

Grok 4.6 measured at all four reasoning efforts on the Artificial Analysis Intelligence Index: low 51.68, medium 59.01, high 60.92 and xhigh 60.01, with the high step the tallest
Four efforts, and the ladder does not climb straight. On the capability index, high outscores xhigh.

The Effort Ladder: More Reasoning Is Not Automatically Better

The short answer: Grok 4.6 exposes four reasoning efforts. Independent measurement scores its default high setting above its top xhigh setting on the capability index, while xhigh scores better on a business-document benchmark. Spending more on reasoning does not reliably buy a better answer.

SpaceXAI's model guide lists the efforts as "Low, medium, high (default), or xhigh" and states that reasoning cannot be disabled — a level must always be selected. That last part matters for anyone porting a workload from a model where thinking is opt-in: there is no zero-reasoning mode to fall back to, so the cheapest possible call on Grok 4.6 is still a reasoning call.

Artificial Analysis has now measured all four. Here is the full ladder from its Intelligence Index v4.1.1, a composite of nine evaluations — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR.

EffortIntelligence Index v4.1.1AA-OmniscienceAA-Briefcase EloMedian output speed (tokens per second)Cost per index task
low51.6825.90131055.43$0.255
medium59.0128.00152257.37$0.783
high (default)60.9230.48156257.75$0.937
xhigh60.0129.32158759.56$1.230

Source: Artificial Analysis leaderboard payload, read August 28, 2026. Index version v4.1.1. Cost per task is what the evaluator measured while running its own battery, not a calculation from list prices.

Read the last two rows against each other, because this is the most actionable single fact on this page. Moving from high to xhigh raises the measured cost per index task by roughly 31 percent and lowers the index score, from 60.92 to 60.01. It also lowers the AA-Omniscience score, from 30.48 to 29.32. What it does raise is the AA-Briefcase Elo, from 1562 to 1587 — a benchmark built around business-document work rather than the nine-evaluation capability composite.

The practical reading is narrow and worth stating carefully: on this instrument, at this reading, the top effort setting is not the best-scoring one. These are single readings, not repeated trials with error bars, so we are not claiming high is definitively superior to xhigh at everything. We are saying that the assumption most teams carry — that the top setting is the one to reach for when accuracy matters — is not supported by the measurement here, and that it costs about a third more. If your workload looks like the index, default to high and treat xhigh as a thing to benchmark on your own tasks rather than a safe upgrade.

At the other end, low is the interesting budget entry: it runs the same battery for about 27 percent of what high costs, at 51.68 against 60.92. That is a real capability drop, not a rounding difference — but for classification, extraction, and routing work, a nine-point index gap on a frontier composite may be irrelevant to your actual task.

What Grok 4.6 Actually Is

The short answer: a text-out frontier model with a 500,000-token window, built by SpaceXAI for long-running agents, served from two US regions through OpenAI-shaped APIs.

SpecificationValue
ReleasedAugust 12, 2026
Context window500,000 tokens
Maximum outputNot published
Input modalitiesText, image
Output modalitiesText only
Reasoning effortslow, medium, high (default), xhigh — cannot be disabled
Function callingYes
Structured outputsYes
Batch APINot supported
API regionsus-east-1, us-west-2
Rate limit at Tier 0150 requests per second, 50 million tokens per minute
Rate limit at Tier 4500 requests per second, 100 million tokens per minute

Sources: SpaceXAI grok-4.6 model page and rate limits documentation, read August 28, 2026.

Access tiers are worth knowing before you plan capacity. SpaceXAI gates rate limits behind five spending tiers unlocked by cumulative spend since January 1, 2026: Tier 0 is free, then $50, $250, $1,000, and $5,000. The documentation states that "once you qualify for a tier, you stay there permanently; tiers never downgrade." For grok-4.6 the ladder runs from 150 requests per second and 50 million tokens per minute at Tier 0 to 500 and 100 million at Tier 4.

One note on the naming, because it will look like a typo otherwise. The company formerly known as xAI completed its rebrand to SpaceXAI earlier in 2026 — we covered it in xAI Is Officially SpaceXAI Now — and the model line kept the Grok name. In practice the two names still coexist across primary sources: Amazon's model card heads the page "xAI — Grok 4.6" while SpaceXAI's own announcement footer reads "© 2026 SpaceXAI LLC." We reproduce whichever name each source uses rather than normalizing them, because which entity a page names is itself information about who wrote it.

Best For

  • Long-context agent loops. A 500,000-token window with function calling and structured outputs, at the lowest input rate among frontier models we track.
  • Cost-sensitive frontier workloads. The measured cost per index task at high sits just under GPT-5.6 Sol and well under Claude Opus 5.
  • Teams already on Cursor or GitHub Copilot. It is in the model picker on both, with no separate contract.
  • Multimodal input pipelines that emit text. Image in, text out — document understanding, screenshot triage, chart reading.

Not For

  • EU data-residency workloads. Covered in detail below — the residency-respecting routing option does not include Europe.
  • Offline batch pipelines. There is no Batch API, and this is not a soft limitation.
  • Anything reading token probabilities. logprobs is accepted and ignored rather than rejected.
  • Image or audio generation. Output is text only. SpaceXAI ships Grok Imagine for visual generation.
Grok 4.6 long-context billing: once a prompt reaches the threshold, every token in the request is repriced rather than only the overflow
Two tiers, one threshold. Crossing it reprices the whole request, not the excess.

What It Costs in Practice

The short answer: the headline rate is $2.00 per million input tokens and $6.00 per million output tokens, but four separate mechanisms can move your actual bill: the long-context threshold, Priority Processing, cache routing, and which host you call through.

1. The long-context threshold

Already covered above and worth repeating because it is the one that surprises people: at 200,000 prompt tokens, every token in the request reprices. Budget models that assume marginal pricing will understate long-context calls by up to a factor of two.

2. Priority Processing doubles the rate

The pricing page lists Priority Processing at 2x the standard rate across all token types. Note that this is a different "twice the price" from the one in the launch announcement, which we untangle in the disagreements section below — one is a service tier, the other is described as a model variant, and they carry the same multiplier.

3. Cache routing is not automatic

This is the most concrete money-saving detail in SpaceXAI's documentation and the easiest to miss. The model guide advises setting a prompt_cache_key parameter on the Responses API, or an x-grok-conv-id header on Chat Completions, to "route a conversation's requests to the same server, making cache hits reliable; without it you often pay full input price on a cache-cold server."

Read that as a price, because it is one. Cached input bills at $0.50 per million tokens against $2.00 uncached — a 75 percent discount that simply does not apply if your requests scatter across servers. For a multi-turn agent replaying a large system prompt on every step, forgetting one header is the difference between the cached rate and the full rate on the bulk of your input tokens.

4. There is a per-request fee for blocked calls

The pricing page lists "a $0.05 usage guideline violation fee per request" for violations caught before generation on the Responses API. Small in isolation, but a misconfigured loop that retries into the same block pays it every time.

5. Bedrock does not charge the list rate

If you route through Amazon, check which inference option you are on before you model costs. The AWS model card prices In-Region and Geo cross-Region inference above the SpaceXAI list rate, and only Global cross-Region inference matches it.

Bedrock inference optionInputOutputCache read
In-Region$2.20 / 1M$6.60 / 1M$0.55 / 1M
Geo CRIS$2.20 / 1M$6.60 / 1M$0.55 / 1M
Global CRIS$2.00 / 1M$6.00 / 1M$0.50 / 1M

Source: Amazon Bedrock model card for Grok 4.6, read August 28, 2026. Standard tier rates.

That is a 10 percent premium for the residency-respecting options, and SpaceXAI's own Bedrock announcement quotes only the Global figures. Bedrock also has its own service tiers with their own multipliers — Priority at 1.75x the standard rate and Flex at 0.5x — which are not the same thing as SpaceXAI's 2x Priority Processing and should not be conflated with it.

Four Grok 4.6 integration limits: logprobs silently ignored, no Batch API, cache routing key required, and reasoning that cannot be disabled
The limits that bite are the quiet ones — accepted and ignored rather than rejected.

Integration Limits Before You Wire It In

The short answer: the constraints that will cost you time are the ones that fail quietly. Grok 4.6 accepts and ignores logprobs, has no Batch API, and behaves differently on Amazon Bedrock than on SpaceXAI's own API.

  • logprobs and top_logprobs fail silently. The models documentation is explicit: they "are not supported by models grok-4.20 and newer. These fields will be silently ignored if set." A rejected parameter shows up in your error handling on the first call. An ignored one does not — confidence scoring, reranking, and any calibration logic downstream simply receives nothing and keeps running as if it worked. If you are porting from a provider that returns them, this breaks silently in production, not loudly in staging.
  • No Batch API. The model page lists it as "Not supported." Overnight bulk classification and offline enrichment jobs built around batch endpoints do not port over; they have to be re-plumbed as rate-limited concurrent calls.
  • Reasoning cannot be turned off. Every call is a reasoning call. There is no non-reasoning mode to fall back to for trivial requests, so the floor on latency and token spend is higher than on models where thinking is opt-in.
  • No published maximum output length. The model page states the context window but no output cap. Plan for the possibility that a long generation is bounded by something you have to discover empirically.

Bedrock behaves differently — read this before assuming parity

Amazon's model card documents a materially different feature set depending on which endpoint you call, and it is not the difference most teams would guess.

Featurebedrock-runtimebedrock-mantle
Structured outputsNot supportedSupported
Server-side tool useNot supportedClient-side tool calling supported
Count tokensNot supported
Intelligent prompt routingNot supported
Application inference profilesNot supported
Prompt cachingSupportedSupported
ReasoningSupportedSupported

Source: Amazon Bedrock model card for Grok 4.6, read August 28, 2026.

Two more Bedrock-specific notes from the same card. The Chat Completions API there "does not return reasoning tokens" — if your accounting or debugging depends on seeing them, use the Responses API and pass include: ["reasoning.encrypted_content"]. And bedrock-runtime does not support in-Region inference for this model at all: requests must name a cross-Region inference profile, either us.xai.grok-4.6 or global.xai.grok-4.6.

Where Grok 4.6 runs: the SpaceXAI API in US regions only, Cursor and Grok Build, GitHub Copilot and Amazon Bedrock, with EU data residency not available
Six surfaces in a fortnight — and the model does not behave identically on all of them.

Where You Can Run It

The short answer: the SpaceXAI API, Cursor, and Grok Build from launch day, then GitHub Copilot, Amazon Bedrock, and Microsoft Foundry within a fortnight. Each host is documented by the vendor, and the feature set is not identical across them.

Date (2026)SurfaceWhat the source says
August 12SpaceXAI API, Cursor, Grok BuildLaunch day; the announcement also names OpenRouter, Vercel, and Cloudflare
August 14GitHub Copilot"now live in GitHub Copilot" across cloud agents, the Copilot CLI, and the VS Code IDE; some businesses and enterprises must enable it in settings
August 18 or 19Amazon BedrockTwo primary sources give two dates — see below
August 26Microsoft FoundrySpaceXAI writes "now available on Microsoft Foundry"

Sources: SpaceXAI announcement posts for the launch, GitHub Copilot, Amazon Bedrock, and Microsoft Foundry, plus the AWS model card, all read August 28, 2026. We have not independently verified the OpenRouter, Vercel, and Cloudflare listings — those are the vendor's claim, not our check. SpaceXAI's Foundry post does not use the words "generally available."

On Cursor specifically, the documentation is more precise than "all plans": Grok 4.6 "is part of the Cursor Models pool on individual and team plans" and is the default speed tier on Pro and higher, while the Start plan — India only — is fixed at medium effort at standard speed. Cursor and SpaceXAI have been part of the same group since the acquisition closed; we covered that in SpaceX Just Bought Cursor for $60B.

The EU question, answered precisely

This deserves its own heading because the honest answer has two halves and most coverage will give you only one.

On SpaceXAI's own API, there is no EU region. The model page lists us-east-1 and us-west-2, and nothing else. We searched the release notes, the model page, the models index, the pricing page, and the rate-limits page for an EU availability statement for Grok 4.6 and did not find one. That absence is notable rather than neutral: the release notes carry an explicit entry announcing EU availability for the previous model, and no equivalent entry exists for this one. We are reporting that we did not find it, which is not the same as reporting that it does not exist.

On Amazon Bedrock, the answer is different and more useful. The AWS model card lists eight EU regions — Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London, and Paris — as able to call Grok 4.6. But every one of them is marked available for Global cross-Region inference only. Geo cross-Region inference, which AWS describes as the option that "routes across Regions within a geography (such as US, EU, and APAC) while respecting data residency," is marked unavailable in all of them; it is available only in the four US regions.

So the precise statement is this: a team in the EU can call Grok 4.6 through Bedrock, but only through the routing mode that AWS defines as going "anywhere worldwide when there are no residency constraints." Access exists. Residency does not. For a regulated European workload, that distinction is the whole answer, and it is the reason "is it available in Europe?" is the wrong question to ask.

Where the Primary Sources Disagree

The short answer: four documented contradictions, all between first-party sources. We report each reading and arbitrate none of them.

1. The knowledge cutoff: February 1 or January 2026

Two pages of the same vendor's documentation give two answers. The models documentation states: "The knowledge cut-off date of Grok 4.6 is February 1, 2026." The Grok 4.6 model guide states "January 2026." These are not different framings of the same date. If your application depends on knowing whether an event from mid-January 2026 is inside the training data, neither page settles it.

2. The default reasoning effort: high or low

SpaceXAI's model guide lists the efforts as "Low, medium, high (default), or xhigh." Amazon's Bedrock model card states the opposite: reasoning effort is configurable through the reasoning parameter as "low" (default), "medium", "high", or "xhigh".

This one has a practical consequence sharper than it first looks, and it is not merely academic given the effort table earlier on this page. If you port code between the SpaceXAI API and Bedrock without setting the effort explicitly, the same request may run at two different efforts on the two surfaces — with a measured index gap of more than nine points between low and high, and roughly a 3.7x difference in cost per task. Set the effort explicitly on every call. Do not inherit a default that two vendors document differently.

3. The Bedrock availability date: August 18 or 19

The AWS model card records a "Model launch date: August 18, 2026." SpaceXAI's own Bedrock announcement is dated August 19, 2026. A one-day gap that matters to nobody operationally, recorded here because it is a clean illustration of how two first-party sources can disagree about a fact neither has any reason to get wrong.

4. The doubled rate card: a context tier, a fast variant, or a service tier

This is the messiest one, and worth untangling carefully because three explanations attach to numbers that are partly identical.

The same figures — $4.00 input, $1.00 cached input, and $12.00 output per million tokens — appear in two different places with two different explanations. SpaceXAI's pricing page presents them as the long-context tier, billed once a prompt reaches 200,000 tokens. Cursor's documentation presents the identical rate card as a product, pricing what it calls "the Fast variant" at $4.00 input, $1.00 cached input, and $12.00 output per million tokens. And the launch announcement says only: "Additionally, there is a fast variant which is twice the price."

Two things complicate this rather than resolving it. First, no grok-4.6-fast model exists in the SpaceXAI price table, which lists exactly seven entries: grok-4.6, grok-4.5, grok-4.3, grok-4.20-0309-reasoning, grok-4.20-0309-non-reasoning, grok-build-0.1, and grok-4.20-multi-agent-0309. Second, the same pricing page independently applies a 2x multiplier to Priority Processing, which is a service tier and not a model at all. So "twice the price" appears in the documentation attached to two genuinely different mechanisms.

We are not going to arbitrate this, and the reason is that all three readings are internally consistent. What a developer needs to take away is operational rather than taxonomic: know which surface you are billed through before you can predict what doubles your bill. Through the SpaceXAI API, a long prompt does it. Through Cursor, a speed choice does it. Through Priority Processing, a service tier does it. The rate card looks the same in the first two cases; the trigger does not.

How Grok 4.6 Stands on Independent Boards

The short answer: strong on one capability index, absent from the agentic-coding board, and provisionally ranked below its own predecessor on human preference. Three instruments, three different answers.

Artificial Analysis Intelligence Index v4.1.1

At high, Grok 4.6 scores 60.92 — level with GPT-5.6 Sol at max effort, which scores 60.93 in the same reading, and behind Claude Opus 5 at 63.05 and Claude Fable 5 at 62.07. On the evaluator's cost-per-task measurement it runs at $0.937 against $0.953 for Sol and $2.337 for Opus 5.

Two cautions on those numbers. The effort levels are not matched across rows — Grok 4.6 is measured at its shipped default while Sol and Opus 5 are measured at max — so this is a comparison of models as configured, not of ceilings. And a cost-per-task figure is what the evaluator measured while running its own battery on real workloads; it cannot be derived from the list prices in the table earlier on this page, nor they from it.

Artificial Analysis Coding Agent Index v1.4

Grok 4.6 is not on it. We read the full board payload on August 28, 2026, counted 57 harness-and-model pairs, and found no Grok 4.6 entry at any effort with any harness. The most recent SpaceXAI pairing is "Grok Build - Grok 4.5 (high)" at 0.6409, which ranks 7th of 57 — a genuinely strong placement for the predecessor.

A method note, because it applies to anyone verifying this themselves: the structured data embedded in that page's markup exposes only nine display rows, and a keyword search over the page will surface "grok-4-6" from the model selector catalog. Neither is the scored board. Concluding either presence or absence from those would be wrong in both directions.

An absence is not a result. It does not mean Grok 4.6 would score badly on agentic coding; it means that as of this reading, no independent agentic-coding measurement of it exists, so anyone citing its agentic standing is citing the vendor. We covered the harness itself in SpaceXAI Open-Sources Grok Build.

Arena text leaderboard

Note first that lmarena.ai now redirects to arena.ai. On the board with a vote cutoff of August 27, 2026, grok-4.6-high sits at rank 48 with a rating of 1461.28 on 3,471 votes, carrying a pre_release flag and a rank interval spanning 25 to 70. Its own predecessor Grok 4.5 sits at rank 36 with 1469.55 on 24,152 votes.

The reserve on that reading is large and we would rather state it than bury it. With 3,471 votes against 24,152, and confidence intervals that overlap between 1464.51 and 1471.39, the board does not cleanly separate the two models. What is harder to wave away is that five SpaceXAI entries currently rank above it, including three from the Grok 4.20 line. On human preference, this flagship has not yet displaced the company's own back catalog.

Alternatives Worth Weighing

  • GPT-5.6 Sol — the closest match on measured capability, at roughly the same cost per index task but double the list input rate. Choose Sol if you need a larger context window or are already inside OpenAI's tooling. Side by side in GPT-5.6 Sol vs Grok 4.5.
  • Claude Opus 5 — about two index points ahead and roughly 2.5x the cost per task. Choose Opus 5 when the top of the capability range is worth paying for, or when you need Claude Code as the harness. See Claude Opus 5 vs Grok 4.5.
  • Kimi K3 — around a point behind on the index at a lower measured cost per task, and it is on the Coding Agent Index where Grok 4.6 is not.
  • Grok 4.5 — the same $2.00 and $6.00 list rates, an established Arena position, and a measured Coding Agent Index placement. Compared against the frontier in Grok 4.5 vs Claude Fable 5.

Who Should Use Grok 4.6

  • Agent developers running long tool-calling loops who can keep prompts under 200,000 tokens and will set a cache routing key.
  • Cost-constrained teams at frontier capability that need index-class reasoning without Opus-class per-task pricing.
  • Cursor and GitHub Copilot users who want frontier reasoning inside an existing subscription rather than a new API contract.
  • Document and screenshot pipelines taking image input and emitting structured text.
  • Teams benchmarking effort settings — the four-rung ladder with published independent scores at every rung makes it unusually easy to pick a tier on evidence.
  • US-based enterprises on Bedrock that can use Geo cross-Region inference and do not need structured outputs on bedrock-runtime.

Frequently Asked Questions

What is Grok 4.6?

Grok 4.6 is SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, released on August 12, 2026. It offers a 500,000-token context window, accepts text and image input and returns text only, and exposes four reasoning efforts — low, medium, high, and xhigh — with reasoning that cannot be disabled. List pricing starts at $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for prompts below 200,000 tokens. We score it 8.0 out of 10.

How much does Grok 4.6 cost?

For prompts below 200,000 tokens, SpaceXAI lists $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens. At or above 200,000 tokens those rates double to $4.00, $1.00, and $12.00. The pricing page states that requests reaching the threshold are billed at the higher rate for all tokens in the request, so crossing it reprices the entire call rather than only the excess. Priority Processing costs 2x the standard rate, and blocked requests carry a $0.05 usage guideline violation fee.

Is xhigh better than high on Grok 4.6?

Not on the capability index. In an Artificial Analysis reading taken on August 28, 2026, Grok 4.6 at high effort scores 60.92 on the Intelligence Index v4.1.1 while xhigh scores 60.01, and high also leads on AA-Omniscience at 30.48 against 29.32. The xhigh setting does score higher on AA-Briefcase, at an Elo of 1587 against 1562, and it costs roughly 31 percent more per index task. High is the shipped default, and on this evidence it is the sensible starting point.

What is Grok 4.6's knowledge cutoff?

Two pages of SpaceXAI's own documentation disagree. The models documentation states that the knowledge cut-off date of Grok 4.6 is February 1, 2026, while the Grok 4.6 model guide states January 2026. We report both rather than arbitrating between them. If your application depends on whether a mid-January 2026 event is inside the training data, neither page settles it.

What is Grok 4.6's default reasoning effort?

It depends on which surface you call. SpaceXAI's model guide lists the efforts as low, medium, high as the default, or xhigh. Amazon's Bedrock model card states that low is the default. Because the measured index gap between low and high exceeds nine points and the cost per task differs by roughly 3.7 times, you should set the effort explicitly on every request rather than relying on a default that two vendors document differently.

Does Grok 4.6 support the Batch API and logprobs?

Neither. The grok-4.6 model page lists Batch API support as "Not supported," which rules out offline bulk workloads built around batching. SpaceXAI's models documentation states that logprobs and top_logprobs are not supported by models grok-4.20 and newer, and that these fields will be silently ignored if set rather than rejected. Because they fail quietly, any downstream confidence-scoring or reranking logic receives nothing and keeps running as though it worked.

Is Grok 4.6 available in the EU?

Partly, and the distinction matters. SpaceXAI's own API serves the model from us-east-1 and us-west-2 only, and we found no EU availability entry for Grok 4.6 in the release notes, model page, models index, pricing page, or rate-limits page. On Amazon Bedrock, eight EU regions including Frankfurt, Ireland, and Paris can call the model — but only through Global cross-Region inference, which AWS defines as routing anywhere worldwide when there are no residency constraints. Geo cross-Region inference, the mode that respects data residency, is available only in the four US regions. Access exists in Europe; data residency does not.

Is there a Grok 4.6 Fast variant, and does it cost double?

The same rate card — $4.00 input, $1.00 cached input, and $12.00 output per million tokens — is described three ways by first-party sources. SpaceXAI's pricing page presents it as the long-context tier billed once a prompt reaches 200,000 tokens. Cursor's documentation presents it as the Fast variant. The launch announcement says only that there is a fast variant which is twice the price. No grok-4.6-fast entry exists in the SpaceXAI price table, which lists seven models. Separately, Priority Processing also carries a 2x multiplier, so "twice the price" attaches to two different mechanisms in the same documentation.

Is Grok 4.6 ranked on the Artificial Analysis Coding Agent Index?

No. On a reading of Coding Agent Index v1.4 taken on August 28, 2026, none of the 57 harness-and-model pairs on the board involves Grok 4.6 at any effort with any harness. The most recent SpaceXAI pairing is Grok Build with Grok 4.5 at high effort, scoring 0.6409 and ranking 7th of 57. Note that a keyword search of that page surfaces Grok 4.6 from the model selector catalog rather than from a scored row, so presence in a page search is not presence on the board.

How does Grok 4.6 compare to GPT-5.6 Sol and Claude Opus 5?

On the Artificial Analysis Intelligence Index v4.1.1 read on August 28, 2026, Grok 4.6 at high effort scores 60.92, GPT-5.6 Sol at max scores 60.93, and Claude Opus 5 at max scores 63.05. Measured cost per index task is $0.937, $0.953, and $2.337 respectively. On list prices, Grok 4.6 charges $2.00 per million input tokens against $4.00 for Sol and $5.00 for Opus 5. The effort levels are not matched across those rows, so this compares the models as configured by the evaluator rather than at their ceilings.

Why does Grok 4.6 rank below Grok 4.5 on Arena?

On the arena.ai text leaderboard with a vote cutoff of August 27, 2026, grok-4.6-high sits at rank 48 with a rating of 1461.28 on 3,471 votes, while Grok 4.5 sits at rank 36 with 1469.55 on 24,152 votes. The newer entry carries a pre_release flag and a rank interval spanning 25 to 70, and the two confidence intervals overlap between 1464.51 and 1471.39. The board does not cleanly separate the two models yet, and lmarena.ai now redirects to arena.ai.

Does Grok 4.6 work the same on Amazon Bedrock as on the SpaceXAI API?

No. Amazon's model card lists structured outputs, server-side tool use, count tokens, intelligent prompt routing, and application inference profiles as unsupported on the bedrock-runtime endpoint, while bedrock-mantle supports structured outputs and client-side tool calling. The Chat Completions API on Bedrock does not return reasoning tokens, so you need the Responses API with reasoning.encrypted_content included to see them. Bedrock also prices In-Region and Geo cross-Region inference at $2.20 per million input tokens against the $2.00 list rate, which only Global cross-Region inference matches.

The Bottom Line

Grok 4.6 earns 8.0 out of 10. It reaches the top group on an independent capability index at the lowest list input rate among the frontier models we track, and the four-rung effort ladder is published, measured, and unusually easy to shop. If your workload is a long-running agent loop with prompts under 200,000 tokens, this is a serious cost argument rather than a marginal one.

What holds it back is not capability, it is coherence. Four separate contradictions sit between first-party sources on this model: two knowledge cutoffs, two default efforts, two Bedrock launch dates, and one rate card explained three ways. Individually they are footnotes. Together they mean that a developer cannot read one page and know what they are buying — the default effort question alone can silently change your cost per task by a factor of nearly four when you port between two surfaces.

The two verified gaps are narrower and cleaner. There is no Batch API, which is a structural exclusion for offline bulk work. And in Europe, access exists through Bedrock but data residency does not — the routing mode that respects it stops at the US border.

Our recommendation: set the reasoning effort explicitly, set a cache routing key, keep prompts below the 200,000-token line, and start at high rather than xhigh. Do those four things and Grok 4.6 is among the best value propositions at the frontier right now. Skip any of them and the bill will not match your model of it.

Sources

Every figure on this page was read from one of these pages on August 28, 2026. No media outlets are used as sources.

Reviewed by Anthony Martinez, ThePlanetTools.ai. Research-led review — we have not run a controlled hands-on evaluation of this model. All specifications, prices, limits, and scores were read from primary sources on August 28, 2026: SpaceXAI developer documentation and announcement posts, the Amazon Bedrock model card, Cursor model documentation, and the leaderboard payloads published by Artificial Analysis and Arena. Benchmark figures are attributed to their evaluators. ThePlanetTools.ai has no affiliation with SpaceXAI, Amazon, Microsoft, GitHub, or Cursor.

Key Features

500,000-token context window
Four reasoning efforts: low, medium, high, and xhigh
Reasoning cannot be disabled
Text and image input, text output
Function calling
Structured outputs
Prompt caching with an explicit routing key
Responses API and Chat Completions API
Priority Processing service tier at twice the standard rate
Five spending tiers with rate limits that never downgrade
Served from us-east-1 and us-west-2
Amazon Bedrock support through bedrock-runtime and bedrock-mantle endpoints

Pros & Cons

Pros

  • Lowest list input rate among the frontier models we track at $2.00 per million input tokens, against $4.00 for GPT-5.6 Sol and $5.00 for Claude Opus 5
  • Measured cost of $0.937 per Artificial Analysis Intelligence Index task at high effort, just under GPT-5.6 Sol at $0.953 and well under Claude Opus 5 at $2.337 (read August 28, 2026)
  • 500,000-token context window with function calling and structured outputs, aimed at long-running agent loops
  • All four reasoning efforts have published independent scores, so the capability-versus-cost tradeoff can be chosen on evidence rather than guesswork
  • Cached input bills at $0.50 per million tokens, a 75 percent discount on the standard input rate
  • Reached six hosting surfaces within a fortnight of launch, including Cursor, GitHub Copilot, Amazon Bedrock, and Microsoft Foundry
  • OpenAI-shaped Responses and Chat Completions APIs, so most existing client code ports with a base URL and model name change

Cons

  • No Batch API, which structurally excludes offline bulk workloads built around batch endpoints
  • logprobs and top_logprobs are accepted and silently ignored rather than rejected, so downstream confidence scoring fails without raising an error
  • First-party sources contradict each other on the knowledge cutoff, the default reasoning effort, and the Bedrock launch date
  • No EU data residency: Amazon Bedrock serves eight EU regions but only through Global cross-Region inference, which AWS defines as routing anywhere worldwide
  • Prompt caching silently degrades to the full input price without an explicit prompt_cache_key or x-grok-conv-id header

Best Use Cases

Long-running agent loops with large tool-calling context
Cost-sensitive frontier reasoning where per-task budget is the binding constraint
Document and screenshot understanding pipelines that emit structured text
In-IDE coding assistance through Cursor or GitHub Copilot
Benchmarking reasoning effort against cost on your own task set
Enterprise deployment on Amazon Bedrock in US regions
Retrieval-augmented workloads sized to stay below the 200,000-token repricing threshold

Platforms & Integrations

Available On

WebREST API

Integrations

SpaceXAI APIOpenAI-compatible SDKAmazon BedrockMicrosoft FoundryCursorGitHub CopilotGrok Build
Anthony M. — Founder & Lead Reviewer
Anthony M.Verified Builder

We're developers and SaaS builders who use these tools daily in production. Every review comes from hands-on experience building real products — DealPropFirm, ThePlanetIndicator, PropFirmsCodes, and many more. We don't just review tools — we build and ship with them every day.

Written and tested by developers who build with these tools daily.

Was this review helpful?

Frequently Asked Questions

What is Grok 4.6?

SpaceXAI's frontier model for long-running agents — 500K context, four reasoning efforts, and the lowest list input rate at the top of the index.

How much does Grok 4.6 cost?

Grok 4.6 costs $2/month.

Is Grok 4.6 free?

No, Grok 4.6 starts at $2/month.

What are the best alternatives to Grok 4.6?

Top-rated alternatives to Grok 4.6 can be found in our WebApplication category, where we've reviewed and scored every tool on ThePlanetTools.ai.

Is Grok 4.6 good for beginners?

Grok 4.6 is rated 7.5/10 for ease of use.

What platforms does Grok 4.6 support?

Grok 4.6 is available on Web, REST API.

Does Grok 4.6 offer a free trial?

No, Grok 4.6 does not offer a free trial.

Is Grok 4.6 worth the price?

Grok 4.6 scores 9/10 for value. We consider it excellent value.

Who should use Grok 4.6?

Grok 4.6 is ideal for: Long-running agent loops with large tool-calling context, Cost-sensitive frontier reasoning where per-task budget is the binding constraint, Document and screenshot understanding pipelines that emit structured text, In-IDE coding assistance through Cursor or GitHub Copilot, Benchmarking reasoning effort against cost on your own task set, Enterprise deployment on Amazon Bedrock in US regions, Retrieval-augmented workloads sized to stay below the 200,000-token repricing threshold.

What are the main limitations of Grok 4.6?

Some limitations of Grok 4.6 include: No Batch API, which structurally excludes offline bulk workloads built around batch endpoints; logprobs and top_logprobs are accepted and silently ignored rather than rejected, so downstream confidence scoring fails without raising an error; First-party sources contradict each other on the knowledge cutoff, the default reasoning effort, and the Bedrock launch date; No EU data residency: Amazon Bedrock serves eight EU regions but only through Global cross-Region inference, which AWS defines as routing anywhere worldwide; Prompt caching silently degrades to the full input price without an explicit prompt_cache_key or x-grok-conv-id header.

Ready to try Grok 4.6?

Get started today

Try Grok 4.6 Now