AI Newsway

Grok 4.7 Holds Its Price and Doubles Its Terminal Score. The Bill Is Still the Argument.

SpaceXAI shipped a new frontier model a month after the last one without moving $2/$6 pricing β€” which reframes every benchmark gap as a cost-per-point question

|4 min read0
AI Summary
SpaceXAI released Grok 4.7 on September 21, 2026 at unchanged pricing of $2 per million input tokens and $6 per million output tokens. The model improves on Grok 4.6 across all seven published benchmarks, nearly doubling Terminal-Bench 4.0 from 20.3% to 38.0%, and beats GPT-5.6 Sol on five. Claude Fable 5.1 still leads multi-hour terminal work by roughly 20 points. Every score comes from SpaceXAI's own testing.
A GPU cluster in a data centre β€” the compute tier behind the extended reinforcement learning run SpaceXAI credits for Grok 4.7's gains
A GPU cluster in a data centre β€” the compute tier behind the extended reinforcement learning run SpaceXAI credits for Grok 4.7's gains

Frontier model launches usually arrive with a price increase attached. Grok 4.7, which SpaceXAI put into general availability on September 21, did not. Input stays at $2 per million tokens and output at $6, the same sticker as Grok 4.6 from a month earlier, and that single decision is what makes the rest of the scorecard interesting.

Key takeaways

  • Grok 4.7 ships at unchanged $2/$6 per million token pricing β€” a fifth of Claude Fable 5.1's input rate and an eighth of its output rate.
  • Its Terminal-Bench 4.0 score climbs from 20.3% to 38.0% in one generation, the largest single jump in the release.
  • Every figure originates with SpaceXAI, and the headline benchmark is maintained by Cursor, a company SpaceX finished buying last month.

Why the flat price matters more than the wins

Buyers of agentic capacity do not purchase benchmark points; they purchase tokens that a long-running task consumes unpredictably. At $2 and $6, Grok 4.7 undercuts GPT-5.6 Sol's $4 and $20 and sits far beneath Fable 5.1's $10 and $50. A team that already budgeted for Grok 4.6 inherits the upgrade at zero incremental cost β€” a rare dynamic in a market where capability jumps normally reprice the tier.

That framing changes how the losses read. Where Grok 4.7 trails, the question becomes whether the gap is worth five to eight times the token bill, not whether the model is behind.

The seven-benchmark table, in one place

SpaceXAI ran Grok 4.7 at xHigh effort against Grok 4.6 at High, GPT-5.6 Sol at Max and Fable 5.1 at Max.

BenchmarkGrok 4.7Grok 4.6GPT-5.6 SolFable 5.1
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11,6571,5461,4871,678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%

*High-effort score.

Counted straight, that is five wins over Sol and three over Fable 5.1, with improvement over Grok 4.6 in every row. The electrical-engineering and legal columns are lopsided enough to suggest domain coverage the rivals simply did not train for. The Terminal-Bench column is the opposite story: Fable 5.1's 57.9% sits almost twenty points clear, and multi-hour terminal work remains the clearest place where paying more still buys something.

What SpaceXAI changed to get there

The company attributes the gains to a larger base model and an extended reinforcement learning run skewed toward tasks measured in hours rather than minutes. Self-verification and long-context handling improved as a result, per the announcement, and the model was additionally trained to operate the Grok Bot harness natively β€” a detail that matters because harness familiarity, not raw reasoning, tends to decide how agents behave on open-ended work.

Safety got a parallel rebuild. SpaceXAI reports 62.4% on LatchBio's biosafety benchmark and a 3.3% pass-through rate for risky dual-use prompts on its in-house HackerBench v0.3, and has begun granting invited security partners access to the model's red-team capabilities.

Two caveats worth carrying forward

The first is provenance. CursorBench 4.0 leads the announcement, and Cursor is now a SpaceX subsidiary following last month's acquisition β€” a vendor-owned benchmark in the headline slot invites independent replication before the number travels. The second is selection: OpenAI's GPT-6 Astra, which launched earlier this month at a premium to Sol, appears only in the secondary charts, where Grok 4.7's 1,695 Elo clears Astra's 1,542 but trails Anthropic's 1,735.

Grok 4.7 is live in Cursor and Grok Build and reachable through the Grok API, third-party harnesses, routers and cloud platforms. As with Grok 4.6's release, the durable test is not the launch table but what neutral evaluators measure over the next few weeks.

FAQ

How much does Grok 4.7 cost?

Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6. A faster variant doubles output speed and doubles the price.

Is Grok 4.7 better than Claude Fable 5.1?

It depends on the task. Grok 4.7 leads on DeepSWE v1.1, EEBench and the Harvey legal benchmark; Fable 5.1 leads on CursorBench 4.0, AA Briefcase, Terminal-Bench 4.0 and HealthBench Professional while costing five times more per input token.

Are the Grok 4.7 benchmark results independently verified?

No. All published scores come from SpaceXAI's own evaluation runs, and CursorBench is maintained by Cursor, which SpaceX acquired last month. Third-party results are not yet available.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Grok 4.6 Reaches the AI Frontier Without Raising Its Price
LLM & Chatbots

Grok 4.6 Reaches the AI Frontier Without Raising Its Price

xAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol while holding pricing flat at $2/$6 per million tokens.

Seung Jung39 days ago
Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing
LLM & Chatbots

Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing

Z.ai confirmed the stealth Ox Alpha model is GLM-5.3-Flash: 320B parameters, 18B active, MIT-licensed weights and sub-dollar output pricing.

Seung Jung25 days ago
GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price
LLM & Chatbots

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

Seung Jung34 days ago
Grok Bot Lets AI Agents Run Their Own Group Chat β€” and Sign Into Your Accounts
LLM & Chatbots

Grok Bot Lets AI Agents Run Their Own Group Chat β€” and Sign Into Your Accounts

SpaceXAI opened a Grok Bot beta where multiple agents coordinate in group chats, assign ownership to each other, and sign into a user's own accounts.

Seung Jung38 days ago
Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price
LLM & Chatbots

Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price

Google's Gemini 3.7 Flash lands three weeks after 3.6 Flash, scoring 65.3% on DeepSWE and shipping at half the price through the end of 2026.

Seung Jung39 days ago
DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing
LLM & Chatbots

DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...

Seung Jung39 days ago