Frontier model launches usually arrive with a price increase attached. Grok 4.7, which SpaceXAI put into general availability on September 21, did not. Input stays at $2 per million tokens and output at $6, the same sticker as Grok 4.6 from a month earlier, and that single decision is what makes the rest of the scorecard interesting.
Key takeaways
- Grok 4.7 ships at unchanged $2/$6 per million token pricing β a fifth of Claude Fable 5.1's input rate and an eighth of its output rate.
- Its Terminal-Bench 4.0 score climbs from 20.3% to 38.0% in one generation, the largest single jump in the release.
- Every figure originates with SpaceXAI, and the headline benchmark is maintained by Cursor, a company SpaceX finished buying last month.
Why the flat price matters more than the wins
Buyers of agentic capacity do not purchase benchmark points; they purchase tokens that a long-running task consumes unpredictably. At $2 and $6, Grok 4.7 undercuts GPT-5.6 Sol's $4 and $20 and sits far beneath Fable 5.1's $10 and $50. A team that already budgeted for Grok 4.6 inherits the upgrade at zero incremental cost β a rare dynamic in a market where capability jumps normally reprice the tier.
That framing changes how the losses read. Where Grok 4.7 trails, the question becomes whether the gap is worth five to eight times the token bill, not whether the model is behind.
The seven-benchmark table, in one place
SpaceXAI ran Grok 4.7 at xHigh effort against Grok 4.6 at High, GPT-5.6 Sol at Max and Fable 5.1 at Max.
| Benchmark | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol | Fable 5.1 |
|---|---|---|---|---|
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
*High-effort score.
Counted straight, that is five wins over Sol and three over Fable 5.1, with improvement over Grok 4.6 in every row. The electrical-engineering and legal columns are lopsided enough to suggest domain coverage the rivals simply did not train for. The Terminal-Bench column is the opposite story: Fable 5.1's 57.9% sits almost twenty points clear, and multi-hour terminal work remains the clearest place where paying more still buys something.
What SpaceXAI changed to get there
The company attributes the gains to a larger base model and an extended reinforcement learning run skewed toward tasks measured in hours rather than minutes. Self-verification and long-context handling improved as a result, per the announcement, and the model was additionally trained to operate the Grok Bot harness natively β a detail that matters because harness familiarity, not raw reasoning, tends to decide how agents behave on open-ended work.
Safety got a parallel rebuild. SpaceXAI reports 62.4% on LatchBio's biosafety benchmark and a 3.3% pass-through rate for risky dual-use prompts on its in-house HackerBench v0.3, and has begun granting invited security partners access to the model's red-team capabilities.
Two caveats worth carrying forward
The first is provenance. CursorBench 4.0 leads the announcement, and Cursor is now a SpaceX subsidiary following last month's acquisition β a vendor-owned benchmark in the headline slot invites independent replication before the number travels. The second is selection: OpenAI's GPT-6 Astra, which launched earlier this month at a premium to Sol, appears only in the secondary charts, where Grok 4.7's 1,695 Elo clears Astra's 1,542 but trails Anthropic's 1,735.
Grok 4.7 is live in Cursor and Grok Build and reachable through the Grok API, third-party harnesses, routers and cloud platforms. As with Grok 4.6's release, the durable test is not the launch table but what neutral evaluators measure over the next few weeks.
FAQ
How much does Grok 4.7 cost?
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6. A faster variant doubles output speed and doubles the price.
Is Grok 4.7 better than Claude Fable 5.1?
It depends on the task. Grok 4.7 leads on DeepSWE v1.1, EEBench and the Harvey legal benchmark; Fable 5.1 leads on CursorBench 4.0, AA Briefcase, Terminal-Bench 4.0 and HealthBench Professional while costing five times more per input token.
Are the Grok 4.7 benchmark results independently verified?
No. All published scores come from SpaceXAI's own evaluation runs, and CursorBench is maintained by Cursor, which SpaceX acquired last month. Third-party results are not yet available.






