SpaceXAI released Grok 4.7 on Monday and, for the third generation running, left the price alone: $2 per million input tokens and $6 per million output tokens, the same rates it charged for Grok 4.5 and Grok 4.6. A fast variant runs at twice the output speed for twice the price. What changed is the scorecard, and the company is leading with a number most model launches bury near the bottom of the table: 19.6 percent on the Harvey Legal Agent Benchmark, which xAI says is nearly three times the 6.7 percent it attributes to Anthropic’s Fable 5.1 Max and almost eight times the 2.5 percent it lists for OpenAI’s GPT-5.6 Sol Max.
“Grok 4.7 is our most capable model for coding and knowledge work,” the company wrote in its announcement. “It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date.” The model is available immediately in Cursor, in Grok Build, through the Grok API, and via third-party coding harnesses, model routers and cloud platforms.
The launch lands just over a month after Grok 4.6 and about two months after Grok 4.5, a cadence that has become the defining feature of xAI’s strategy since SpaceX absorbed the lab. Every score the company published is self-reported, and the comparison set is notable for who is missing: OpenAI’s newer and more expensive GPT-6 Astra shows up only in a single GDPval chart, not in the main benchmark table.
The numbers
On xAI’s own table, Grok 4.7 at xHigh effort scores 46.3 percent on CursorBench 4.0, up from 40.4 percent for Grok 4.6 and ahead of the 41.7 percent listed for GPT-5.6 Sol, though behind Fable 5.1 at 51.8 percent. On DeepSWE v1.1 it reaches 71.0 percent at high effort, against 65.2 percent for its predecessor, 72.7 percent for GPT-5.6 Sol and 70.0 percent for Fable 5.1. Terminal-Bench 4.0 jumps to 38.0 percent from 20.3 percent, the largest single-version gain in the table, but still trails Fable 5.1 by almost 20 points. HealthBench Professional comes in at 56.7 percent, behind both GPT-5.6 Sol at 60.5 percent and Fable 5.1 at 62.1 percent. On EEBench, an electrical engineering test, Grok 4.7 posts 64.0 percent and leads the field. On GDPval, the professional-tasks benchmark, it scores 1,695 Elo, up 90 points from Grok 4.6 and behind Fable 5.1’s 1,735 but ahead of the 1,542 xAI lists for GPT-6 Astra.
The technical story behind those gains, per xAI, is a new and larger base model, a longer reinforcement-learning run weighted toward tasks that take many hours to complete, and training the model to natively understand the Grok Bot harness. The company also says it rebuilt its safeguard stack from scratch, calling Grok 4.7 “the strongest model we’ve tested on refusals and jailbreak resistance.” It reports 62.4 percent on LatchBio’s biosafety benchmark and says the model lets through only 3.3 percent of risky dual-use prompts on HackerBench v0.3, its internal cyber-misuse test.
Independent evaluators are less impressed. The Decoder reported that on the Artificial Analysis Intelligence Index v4.3.2, which aggregates ten benchmarks, Grok 4.7 scores 46 and lands mid-pack, while Claude Fable 5.1 and GPT-6 both score 53. Artificial Analysis’s own Terminal-Bench 4.0 run puts Grok 4.7 at 26 percent, well under the 38 percent xAI reported and far behind GPT-6 Astra at 60 percent and Fable 5.1 at 55 percent.
Why it matters
The Grok 4.7 release is a pricing argument dressed up as a capability announcement. At $2 and $6, xAI is charging half of GPT-5.6 Sol’s $4 input rate and a fifth of Fable 5.1’s $10, and on output the gap is starker: $6 against $20 and $50. The company’s own CursorBench cost-per-task chart makes the pitch explicitly, showing Grok 4.7 reaching roughly 46 percent at about $6 per task while Fable 5.1 needs close to $17 to hit 51.8 percent. For teams routing high volumes of coding traffic, that is the relevant math, and it is why xAI keeps freezing the price while pushing the benchmarks.
But the freeze cuts both ways, because the price floor is now being attacked from below. Xiaomi published MiMo-V2.6 under an MIT license and reports 72.57 on DeepSWE, a hair above Grok 4.7, though that figure comes from Xiaomi’s own harness and has not been reproduced on the public leaderboard. StepFun’s Step 5 Preview, released a day earlier, charges $1 input and $2.70 output and posts 67.7 on DeepSWE, with open weights promised for mid-October. Grok 4.7 is cheap relative to San Francisco and expensive relative to Beijing, and it is closed-weight either way.
The Harvey legal number is the most interesting claim and the hardest to interpret. A 19.6 percent score is not a passing grade for legal agent work, but if the comparison holds it is a large relative lead in a domain xAI is actively selling into through its Legal solutions page. Harvey’s benchmark is young and its leaderboard sparse, so the gap could reflect harness differences as much as capability. The same caution applies to CursorBench, which is built by Cursor, a company SpaceX finished acquiring last month. When the lab and the benchmark share a parent, the headline number deserves a second look.
What to watch
The first thing to watch is whether third parties can reproduce the Harvey and Terminal-Bench figures, given that Artificial Analysis already has Grok 4.7 12 points below xAI’s own Terminal-Bench score. The second is whether Cursor, now a sibling company, starts defaulting users to Grok 4.7 in ways that shift real-world share regardless of what the composite indexes say. And the third is the cadence: three Grok releases in roughly ten weeks at a fixed price suggests xAI is optimizing for momentum over margin, and the next test is whether the Chinese open-weight labs, now shipping models at a fifth of Grok’s price, force the $2 and $6 line to finally move.
“It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date.”— SpaceXAI, Grok 4.7 launch announcement