SPAWNSY

Grok 4.7 Arrives a Month Late: Better Than 4.6 Everywhere, Still Behind Fable 5.1

xAI shipped Grok 4.7 31 days after Musk's original deadline: same price, and a Terminal-Bench score almost double that of 4.6. xAI's own table still shows Fable 5.1 winning four of seven tests, and the independent index gives Grok 46 points against 53.

AuthorTwenZySPAWNSY Editorial Desk
PublishedSeptember 21, 2026
Read time6 min
SectionTech
Views546
Share
Grok 4.7 Arrives a Month Late: Better Than 4.6 Everywhere, Still Behind Fable 5.1

Grok 4.7 shipped on September 21, 31 days later than Elon Musk had promised, and xAI calls it a "notable improvement" over Grok 4.6 at the same price and speed. The company's numbers back that up: Terminal-Bench went from 20.3% to 38.0%, and the new model beats its predecessor on every one of the seven tests in xAI's table. That same table also shows what the announcement leaves out. The model Musk described on September 14 as "roughly on par with Opus 5.0, not 5.1" still loses to Claude Fable 5.1 on four of the seven tests, and Artificial Analysis's independent index puts it at 46 points against 53 for Fable and GPT-6.

What xAI actually shipped

According to xAI (which now signs its announcements as SpaceXAI), Grok 4.7 is a larger base model than Grok 4.6, further trained with extended reinforcement learning on harder tasks, with an emphasis on verifying its own output. Pricing is unchanged: $2 per million input tokens and $6 per million output tokens. A fast variant costs twice as much and generates twice as quickly. The model is available in the xAI API, in Cursor, in Grok Build, on third-party coding platforms and through model routers and cloud services.

xAI did not disclose a parameter count. Musk talked about a 2.1 trillion parameter model in July, but nothing official confirms it. Two days after saying on September 11 that 4.7 "needs a few more days," he announced Grok 4.8 at 2.5 trillion parameters with no release date. A week later 4.7 shipped anyway.

xAI's table: how much improvement is really there

The improvement over Grok 4.6 is real. Engineering reasoning and terminal work gained the most, and all seven entries went up. The table needs a careful read, though, because xAI sets Grok 4.7 at its highest effort setting (xHigh) against Grok 4.6 at High. Part of the jump may come from the new model getting more compute per answer.

* The DeepSWE score for Grok 4.7 is, per xAI, a High-effort result. All figures come from xAI's own materials.

Against GPT-5.6 Sol Max, the model xAI used to measure its predecessor, Grok 4.7 wins five of seven tests, though in Terminal-Bench the margin is 0.7 points. Fable 5.1 is clearly stronger where agentic terminal work is concerned (57.9% against 38.0%) and in medical reasoning. Grok beats it in electrical engineering and legal tasks, and xAI already pushed the second niche with Grok 4.6, as we wrote in our article on that launch.

One note about the table itself: it opens with CursorBench, an internal test suite from Cursor, the company SpaceX bought in August, which we covered in the story on OpenAI's clash with Cursor. Grok 4.7 loses to Fable on it, so it doesn't look like a test picked to flatter the result, but it is worth knowing who writes it.

The independent measurement is less flattering

On the Artificial Analysis Intelligence Index (v4.3.2), Grok 4.7 scores 46, while Claude Fable 5.1 and GPT-6 score 53 each. The Decoder, which reported the gap, notes that it widens in agentic coding. On Terminal-Bench 4.0, the figures that outlet cites are 26% for Grok 4.7, 55% for Fable 5.1 and 60% for GPT-6 Astra. That number differs from the 38.0% in xAI's table, probably because of different settings and effort modes, and it shows how much a score depends on how it is measured.

Neither GPT-6 nor Astra appears in xAI's table, even though OpenAI released it on September 3, as we covered in our piece on GPT-6 Astra. It is OpenAI's strongest model. The table's numbers match xAI's official page, but limiting the comparison to Sol Max is convenient for the company.

Price is the one argument that needs no trust

xAI's comparison yields simple arithmetic. Grok 4.7 costs $2 and $6 per million tokens, GPT-5.6 Sol Max $4 and $20, and Fable 5.1 Max $10 and $50. Against Fable that is one fifth of the input price and more than eight times cheaper on output. Anyone running a model thousands of times a day in an agent or a document pipeline pays real money for such gaps, which is why a model a few points weaker can win the contract.

There is a limit to that. A token price is not a task price: if the weaker model needs more attempts or more reasoning tokens, the savings shrink. xAI doesn't show that, and independent measurements of whole-task cost for Grok 4.7 are still to come.

The story of the deadlines

On July 24 Musk said Grok 4.6 would arrive in two weeks and Grok 4.7 in four, meaning around August 21. Version 4.6 shipped on August 12 and 4.7 only on September 21. In early September Musk talked about roughly ten days. On September 11 he said the model "needs a few more days," explaining that he may have applied too large a penalty for response length in training, which made the model "give up too early" on tasks it could solve. On September 14 he judged that 4.7 would be "roughly on par with Opus 5.0, not 5.1, better in some ways, worse in others," and said multimodal work needed fixing.

That September 14 sentence is the best yardstick for the launch. An independent score of 46 against 53 for the top models matches how Musk himself described his model a week earlier.

Safety: xAI's claims

xAI says Grok 4.7 is "the strongest model we've tested" on refusals and jailbreak resistance, with 62.4% on the LatchBio biosafety test and 3.3% of risky prompts passed through on HackerBench v0.3. These are the company's own figures, without independent verification, so we treat them as a promise to check.

Grok 4.7 is a clear step forward from 4.6 and nothing more. The model nearly doubled its terminal score, improved electrical engineering and legal tasks, and the price stayed the same. Anyone who built on Grok 4.6 has a good reason to upgrade.

The frontier is still Fable 5.1 and GPT-6, though. The independent ranking gives Grok 4.7 46 points against 53, the gap in agentic coding is large, and xAI's table skips OpenAI's newest model. After a month of delay Musk himself rated the model at Opus 5.0 level. In our view the launch reads like the closing of a long-dragged promise: the model is better than 4.6, but not by enough to justify that many months of announcements.

It does have a price argument its rivals cannot match: an input price five times lower than Fable's. On the tasks where Grok wins (engineering, legal documents, part of coding), that is enough to pick it. For agentic terminal work and tasks where absolute quality counts, the lead of Anthropic and OpenAI is too large.

Recommendation: test Grok 4.7 on your own tasks and your own token bill instead of trusting one table. If cost is your main constraint, it is a model worth trusting. If you want the best result at any price, you buy elsewhere.

Comments

Discussion

Join the conversation around this story.

0 entries

Join the discussion

Sign in to comment and reply to other readers.

Sign in

No comments yet

Start the discussion first.

Read next

All posts
Tech

Qwen3.8 Max: Alibaba answers a rival it partly owns

Two days after Kimi K3 shook the chip market, Alibaba answered with its own reveal: Qwen3.8 Max, a 2.4 trillion-parameter model. The catch: Alibaba owns roughly 36 percent of Moonshot AI, the company behind Kimi K3.

Flavi7/21/20263 min