The Jev model is now integrated and live · Welcome to try it
Blog
Guide

Grok 4.7 Review

Grok 4.7 is a credible same-price upgrade over 4.6 for long coding and knowledge-work agents, but xAI's launch scores are not a substitute for a matched workload test.

13 min readOmniaKey
Grok 4.7Grok 4.6API pricingcoding benchmarkmodel review

This Grok 4.7 review reaches a narrower conclusion than the launch headlines. Grok 4.7 is the sensible model to evaluate first for long-horizon coding, tool use, and professional knowledge work because it improves every Grok 4.6 row in xAI's published comparison while keeping the same standard API price and 500K context window. It is not a proven universal winner: the launch table mixes reasoning settings, and we did not run a billable first-hand benchmark.

SpaceXAI released grok-4.7 on September 21, 2026. The public API is live, but OmniaKey did not list the model when checked on September 22. Until the live catalog changes, use Grok 4.6 on OmniaKey or call 4.7 through a provider that actually exposes the exact ID.

Evidence checked September 22, 2026. Specifications and launch scores below come from SpaceXAI's announcement and developer documentation. Artificial Analysis had published an early xhigh Intelligence Index result, but its speed and provider-price fields were still incomplete. We did not use a private preview, xAI credential, or first-hand latency run. Vendor scores, independent results, and our recommendation are labeled separately.

Grok 4.7 review: the verdict

Test 4.7 before 4.6 for new long-running agent work. The strongest release evidence is the generation-over-generation change: xAI reports higher results for 4.7 on all seven rows in its direct table, including large gains on Terminal-Bench 4.0 and EEBench, without raising standard token rates.

Do not migrate on the headline alone. xAI compares Grok 4.7 at xhigh, Grok 4.6 at high, and the other frontier models at max; one DeepSWE result is separately marked high effort. Those settings can change token use, latency, and task behavior.

Keep 4.6 as a rollback until your own tasks pass. Both generations expose 500K context and the same direct rates. A staged evaluation costs less than discovering a tool, refusal, verbosity, or output-format regression in production.

Release facts at a glance

ItemGrok 4.7 contract
Release dateSeptember 21, 2026
API model IDgrok-4.7
Input / outputText and image / text
Context window500,000 tokens
Knowledge cutoffMay 2026
Reasoning effortlow, medium, high (default), xhigh
APIsResponses API and Chat Completions
CapabilitiesFunction calling, structured outputs, web search, X search, code execution
Batch APINot supported
Direct global price below 200K prompt tokens$2 input / $0.50 cached input / $6 output per 1M
Direct global price at or above 200K prompt tokens$4 input / $1 cached input / $12 output per 1M

The 500K figure is capacity, not a free allowance. Once the prompt reaches 200K tokens, xAI says the higher rates apply to all tokens in that request. The start guide lists no text output limit, but account rate limits, latency, budget, and endpoint behavior still bound real workloads.

What changed from Grok 4.6?

SpaceXAI says Grok 4.7 uses a new, larger base model, a longer reinforcement-learning run, and a harder training mix weighted toward tasks that take hours. It also claims better self-verification, longer-context management, native understanding of the Grok Bot harness, and a new safeguard stack.

Those are useful release details, not independently audited architecture results. The company does not publish a parameter count in the launch post or current model contract.

Decision pointGrok 4.7Grok 4.6What it means
Base/trainingNew larger base; longer RL on harder long tasksPrevious generationPlausible reason to retest long agents
Direct standard price$2 / $0.50 / $6$2 / $0.50 / $6No list-price penalty for upgrading
Long-context price$4 / $1 / $12$4 / $1 / $12Same 200K threshold economics
Context500K500KNo capacity upgrade
Reasoninglow through xhigh; high defaultlow through xhigh; high defaultPin the same effort in a comparison
BatchNot supportedNot supportedDo not budget an assumed Batch discount
Launch emphasisCoding, agents, knowledge work, self-checkingCoding and long-running agents4.7 targets reliability over a new modality

The practical upgrade is claimed task performance at equal list price, not a larger context window or cheaper token. The Grok 4.6 API pricing guide remains the detailed budget reference for the route currently listed by OmniaKey.

Grok 4.7 benchmarks: what the launch table says

These are SpaceXAI-reported results. They were not reproduced by OmniaKey. Settings shown by the publisher are 4.7 xhigh, 4.6 high, GPT-5.6 Sol max, and Fable 5.1 max; the asterisk marks 4.7 DeepSWE at high effort.

EvaluationGrok 4.7Grok 4.6GPT-5.6 SolFable 5.14.7 vs 4.6
CursorBench 4.046.3%40.4%41.7%51.8%+5.9 pp
DeepSWE v1.171.0%*65.2%72.7%70.0%+5.8 pp
EEBench64.0%53.0%39.4%56.4%+11.0 pp
AA Briefcase v1.11,6571,5461,4871,678+111 Elo
Terminal-Bench 4.038.0%20.3%37.3%57.9%+17.7 pp
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%+3.8 pp
HealthBench Professional56.7%48.5%60.5%62.1%+8.2 pp

The table supports one solid claim: on xAI's own evaluation setup, 4.7 improves over 4.6 across every displayed task family. It does not support “Grok 4.7 beats every model.” Fable leads CursorBench, AA Briefcase, Terminal-Bench, and HealthBench in the displayed table; GPT-5.6 Sol leads DeepSWE. Price, reasoning effort, harness, token use, and acceptance criteria also differ.

For coding, the Terminal-Bench gain is the most interesting signal because it targets multi-hour terminal work. It is also where Fable's reported lead is largest. Run both the success criterion and the failure categories, not just the aggregate percentage.

What independent evidence exists?

Artificial Analysis had added Grok 4.7 at xhigh by our fact-check. Its model page showed an Intelligence Index of 46 and described the model as among the leading, well-priced systems in its comparison. It also showed high output-token use: 240M tokens across its index versus a 92M median on the page.

That snapshot is useful but early. Output speed was still N/A, and the page's provider-price fields were not populated reliably. We therefore do not use it to validate xAI's “same speed” statement, the Fast variant, or a cost-per-task figure. The score can also move as the evaluation set or run completes.

The right reading is:

  • official evidence shows a consistent 4.6-to-4.7 improvement;
  • one independent suite already places xhigh in the frontier range;
  • no stable independent speed, p95 latency, or matched 4.6 cost comparison was available at check time;
  • your workload remains the deciding test.

Grok 4.7 API pricing and long context

Grok 4.7 keeps the same direct global rate card as Grok 4.6. The threshold changed in wording from many older summaries: the current model page says below 200K versus at or above 200K, not “200K is still standard.”

RoutePrompt rangeInputCached inputOutput
xAI globalBelow 200K$2.00$0.50$6.00
xAI globalAt or above 200K$4.00$1.00$12.00
xAI US regionalBelow 200K$2.20$0.55$6.60
xAI US regionalAt or above 200K$4.40$1.10$13.20

All values are USD per one million tokens. The US endpoint keeps inference in the United States and carries a 10% token-price premium. Tool calls, storage, taxes, currency conversion, and a third-party platform's own quote are separate.

Three simple global-endpoint examples show why the threshold matters:

WorkloadCalculationModel-token cost
100K uncached input + 20K output0.10 x $2 + 0.02 x $6$0.32
250K uncached input + 20K output0.25 x $4 + 0.02 x $12$1.24
50K uncached + 200K cached + 20K output0.05 x $4 + 0.20 x $1 + 0.02 x $12$0.64

Cached pricing applies only when the provider reports a cache hit. xAI recommends a prompt_cache_key on Responses or x-grok-conv-id on Chat Completions so related turns reach the same server. Context compaction matters for long tool loops because a repeated 200K-plus prompt doubles all three token rates.

The Fast variant needs a careful reading

Grok 4.7 Fast is the same model on faster infrastructure. It is available only in Cursor and Grok Build, not on the public xAI API or Grok Build's free tier.

The current pricing page describes Fast as “twice the standard token rates,” then lists $4 / $1 / $12 below 200K and $6 / $1.50 / $18 above 200K. The second row is 1.5x the public long-context row, not 2x. Quote the table rather than deriving a rate from the prose, and recheck the live bill before a large run. Cursor charges its variant through the Cursor plan.

“Twice the output speed” is also a provider statement, not a latency guarantee. Decode throughput does not include queue time, prompt prefill, time to first token, tools, network delay, or retries.

API details that matter in an agent

The Responses API always returns reasoning.encrypted_content for grok-4.7, even when include does not request it. Pass those reasoning items back unchanged on the next turn so a multi-turn agent preserves the model's reasoning state. Chat Completions behavior is unchanged.

Before migration, verify:

  1. the exact returned model ID and endpoint;
  2. the same reasoning_effort on both generations;
  3. tool-call schema validity and retry behavior;
  4. prompt, cache-read, reasoning, and visible output tokens;
  5. time to first token, end-to-end latency, and accepted-task cost;
  6. output length, because an xhigh result can spend more tokens even at the same rate;
  7. rollback to grok-4.6 without changing the rest of the harness.

For a direct xAI Responses request:

bash
curl https://api.x.ai/v1/responses \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.7",
    "prompt_cache_key": "repo-review-v1",
    "input": "Review this patch, run the acceptance criteria, and report any remaining failure."
  }'

Use a non-secret stable cache key, not an API key, email address, repository secret, or customer identifier. The direct xAI base URL and credential above are not OmniaKey settings.

Where Grok 4.7 is available

SurfaceStatus on September 22, 2026Important boundary
xAI APIAvailable as grok-4.7Standard variant; Fast is not public API
Grok BuildDefault modelFast is paid and excluded from the free tier
CursorAvailable on all plans per xAIFast billing follows Cursor's plan
GitHub CopilotGradual rollout announcedPro, Pro+, Max, Business, Enterprise; admin policy can disable it
OpenRouter, Vercel, CloudflareListed by xAI as gatewaysEach owns its route, price, and availability
OmniaKeyNot listed at fact-check timeCheck the live model catalog; do not guess a route

GitHub says the rollout covers VS Code, Visual Studio, Copilot CLI, the cloud agent, the Copilot app, JetBrains, Xcode, and Eclipse. “Rolling out” does not mean every account sees the picker immediately.

A reproducible Grok 4.6 to 4.7 test

Use 20 to 50 tasks drawn from your real failure distribution rather than a collection of easy prompts. Freeze the repository revision, system prompt, tools, permissions, effort setting, timeout, and acceptance commands.

Record at least:

  • accepted without repair, accepted after retry, and failed;
  • tool-schema errors, incorrect edits, test failures, refusals, and loops;
  • input, cached input, reasoning, and output tokens;
  • first-token and end-to-end latency at p50 and p95;
  • provider-reported model ID and total billed amount;
  • human review or repair minutes.

Choose 4.7 when it improves accepted-task cost or reduces meaningful failure modes. Keep 4.6 when results are equivalent and your existing route is operationally simpler. For cross-family routing, use the coding-agent model guide rather than generalizing from one launch table.

Common Grok 4.7 review mistakes

Calling vendor scores an independent review

The seven-row launch table is useful, but SpaceXAI selected and ran it. Attribute it and keep the effort settings visible.

Claiming a 2.1T parameter count

Parameter-count discussion appeared before release, but the launch post and current developer contract do not publish a number. “New, larger base model” is the confirmed wording.

Treating 500K context as free

It is the maximum prompt capacity. At 200K tokens, the whole request enters the higher direct-price tier.

Assuming Fast is a public API model ID

It is a serving option in Cursor and Grok Build. The public xAI API exposes grok-4.7, not a documented Fast slug.

Assuming release means OmniaKey availability

Provider launch and gateway onboarding are separate events. There was no public OmniaKey 4.7 route on the fact-check date.

Frequently asked questions

When was Grok 4.7 released?

SpaceXAI released Grok 4.7 on September 21, 2026. The direct API model ID is grok-4.7.

Is Grok 4.7 better than Grok 4.6?

It is the stronger first candidate. xAI reports higher 4.7 results on every displayed 4.6 comparison at the same direct token price. A matched workload test is still required because effort, token use, tools, and failure modes affect production value.

How much does the Grok 4.7 API cost?

On the global xAI endpoint, prompts below 200K tokens cost $2 input, $0.50 cached input, and $6 output per million tokens. At or above 200K, the rates are $4, $1, and $12.

What is the Grok 4.7 context window?

The documented context window is 500,000 tokens. The current start guide also says there is no text output limit, but rate, time, and budget limits still apply.

How many parameters does Grok 4.7 have?

SpaceXAI has not published a parameter count in the launch post or developer model page. It only confirms that 4.7 uses a new, larger base model than 4.6.

Does Grok 4.7 support images and tools?

Yes. It accepts text and image input, returns text, and documents function calling, structured output, web search, X search, and code execution.

Is Grok 4.7 available on OmniaKey?

Not at the September 22 fact-check. The public catalog listed Grok 4.6 and 4.5 but not grok-4.7. Check the live catalog rather than assuming availability from this review.

Sources checked

Fact-checked September 22, 2026. Model availability, dynamic independent scores, prices, and platform rollouts can change. Recheck the primary sources and the live catalog before a production migration.