Grok 4.7 Launches at $2 Per Million Tokens, Targets Coding and Legal AI Work

N
Navs
Published on September 23, 20266 min read
Grok 4.7 Launches at $2 Per Million Tokens, Targets Coding and Legal AI Work

xAI released Grok 4.7 on September 21, 2026, pricing the frontier model at $2 per million input tokens and $6 per million output tokens — the same as Grok 4.6, and a fraction of what competitors charge for comparable models.

The model ships with a 500,000-token context window, accepts text and image inputs, and outputs text only. xAI describes it as "SpaceXAI's most powerful model for coding and knowledge work," trained with a longer reinforcement-learning run on harder tasks, including problems that take many hours to complete

Grok 4.7 is available now on the xAI API, in Cursor, in Grok Build, on OpenRouter, on Vercel's AI Gateway, and is rolling out in GitHub Copilot. A faster variant, Grok 4.7 Fast, runs at twice the token rates and is available only through Cursor and Grok Build .

Pricing and the Frontier Model Gap

The launch subhead reads "Twice as fast, at half the price of comparable models." That comparison is not against Grok 4.6 — the per-token prices are identical — but against the frontier models xAI chose for its benchmark table:

ModelInput (per 1M tokens)Output (per 1M tokens)
Grok 4.7$2$6
GPT-5.6 Sol$4$20
Fable 5.1$10$50

Grok 4.7 costs roughly one-eighth of Fable 5.1's output price and one-third of GPT-5.6 Sol's. For prompts above 200,000 tokens, Grok 4.7's pricing doubles to $4 input and $12 output per million tokens .

That pricing pressure matters for the broader market. If developers can route coding and knowledge-work tasks to a model at $2/$6 without large quality sacrifices, the economics of running AI agents for hours-long workflows shift. Frontier models from OpenAI and Anthropic now compete not just on capability but on cost-per-task. For a developer running a multi-hour coding agent that generates 500 million output tokens over a week, the difference between Grok 4.7 and Fable 5.1 is $3,000 versus $25,000 — a gap that reshapes which models are viable for sustained agentic workflows, not just one-off queries.

Benchmark Performance Against GPT-5.6 Sol and Fable 5.1

xAI published its own benchmark table comparing Grok 4.7 at xHigh reasoning effort against Grok 4.6, GPT-5.6 Sol, and Fable 5.1. All figures below are from xAI's own testing .

BenchmarkGrok 4.7Grok 4.6GPT-5.6 SolFable 5.1
CursorBench 4.0 (software engineering)46.3%40.4%41.7%51.8%
DeepSWE v1.1 (software engineering)71.0%65.2%72.7%70.0%
Terminal-Bench 4.0 (multi-hour terminal work)38.0%20.3%37.3%57.9%
EEBench (electrical engineering)64.0%53.0%39.4%56.4%
AA Briefcase v1.1 (multi-hour office work)1,6571,5461,4871,678
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional (clinical reasoning)56.7%48.5%60.5%62.1%

Grok 4.7 beat Grok 4.6 on all seven benchmarks, with the largest gain on Terminal-Bench 4.0 — nearly doubling from 20.3% to 38.0%. Against GPT-5.6 Sol, Grok 4.7 led on four of seven benchmarks, including a striking gap on the Harvey Legal Agent Benchmark (19.6% vs. 2.5%). Against Fable 5.1, the picture is mixed: Grok 4.7 led on EEBench and Harvey Legal, but trailed on CursorBench, Terminal-Bench, HealthBench, and GDPval .

The Harvey Legal result stands out. Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark — nearly eight times GPT-5.6 Sol's 2.5% and nearly three times Fable 5.1's 6.7%. Legal AI is a high-value market where accuracy and reliability matter more than raw speed, and where providers like Harvey — reportedly valued at $15.6 billion — have built their businesses on top of frontier models from OpenAI and Anthropic.

The Token Usage Catch

Artificial Analysis published its independent evaluation 21 minutes after xAI's announcement. Grok 4.7 scored 46 on the Artificial Analysis Intelligence Index, ranking 16th of 655 models tested. But the gains came with a cost: Grok 4.7 used 81,000 output tokens per Intelligence Index task, more than double Grok 4.6's 36,000 tokens.

At the same per-token price as Grok 4.6, a task that previously cost one unit would cost approximately two units with Grok 4.7. The model's reasoning improvements come from working longer on difficult problems and checking its own work more carefully — which means more tokens generated before the final answer.

For comparison, GPT-6 Astra used 27,000 output tokens per task in the same comparison, and Muse Spark 1.3 Max used 60,000. Grok 4.7's hallucination rate dropped from 34% (Grok 4.6) to 29%, and its accuracy held steady at 47% vs. 48%.

The practical consequence: Grok 4.7's headline price advantage narrows at the task level. Developers evaluating cost-per-task rather than cost-per-token should factor in the roughly 2x token consumption before switching.

Safety Numbers, Published for the First Time

xAI published the first hard safety figures for a Grok launch:

Safety BenchmarkGrok 4.7 Result
LatchBio biosafety benchmark62.4%
HackerBench v0.3 risky dual-use prompts allowed through3.3%

xAI also offered invite-only red-team access to cybersecurity partners ahead of launch. The company describes Grok 4.7 as having its "best-calibrated safeguards to date," and the model was trained to natively understand the Grok Bot harness — xAI's agentic framework for multi-step tasks .

These numbers give developers something concrete to evaluate — a departure from xAI's previous launches, which included no published safety metrics. Whether 3.3% of risky dual-use prompts passing through is acceptable depends on the use case, but it is at least measurable.

Musk's Claims Versus Reality

Elon Musk made several pre-launch claims about Grok 4.7. Here is how they held up:

  • "2.1 trillion parameters" (July 28) — xAI published no parameter count. Unconfirmed.
  • "SpaceX engineering data in supplemental training" (July 21, August 12) — the launch post does not mention SpaceX data. Absent.
  • "Will exceed all current models" (August 12) — Fable 5.1 Max led on four of seven benchmark rows and GDPval. Not met.
  • "Roughly on par with Opus 5.0, not 5.1" (September 14) — Opus does not appear in the benchmark table. Ungradable.
  • "Multimodal performance needs fixing" (September 14) — no multimodal benchmark row was included. Not addressed.

The model also shipped 31 days after Musk's first stated target of "about August 21" and 10 days after his September 11 statement that it needed "a few more days to cook". The release-history table shows accelerating cadence at xAI: Grok 4 shipped July 2025, Grok 4.6 arrived August 12, 2026, and Grok 4.7 followed just 40 days later.

Vercel offered a 40% discount on Grok 4.7 for one week beginning September 21, and OpenRouter listed the model within two minutes of the announcement. GitHub's changelog described the model as "designed for agentic coding and complex, multistep workflows," though Copilot had not yet published the premium-request multiplier for the model.

What Comes Next

Grok 4.7 lands in an increasingly crowded frontier-model market where GPT-6 Astra (released September 3 by OpenAI), Fable 5.1, and Grok 4.7 now compete on both capability and price. The next question is whether xAI can close the gap on CursorBench and Terminal-Bench — the two coding benchmarks where Fable 5.1 leads by significant margins — without further inflating token usage. Grok 4.6 shipped 27 days before Grok 4.7; if xAI maintains that cadence, Grok 4.8 could arrive before November.


Sources: Introducing Grok 4.7, xAI Docs, CellCog, AIToolsRecap, HeadsUpAI

Share this article