AI & ML 4 min read

xAI's Grok 4.7 Leads Two of Seven Benchmarks on Its Own Scorecard

Fable 5.1 takes four and GPT-5.6 Sol one. The wins are electrical engineering and a legal agent test, where the margin is threefold.

Maya Lin
Digital Platforms Analyst
Published 25 Sep 2026, 11:01 PM (SGT)
Share:
Charts and graphs displayed on a computer monitor Charts and graphs displayed on a computer monitor Photo by yatsusimnetcojp on Pixabay
Advertisement

25 SEP 2026 — xAI released Grok 4.7 on 21 September and published a comparison against three rivals across seven benchmarks. The first thing to know about the table is that every figure in it was produced by the company that sells the model.

On its own scorecard, the new model leads two of the seven.

What the table actually shows

The published comparison sets Grok 4.7 against its predecessor, against GPT-5.6 Sol and against Fable 5.1. Fable leads four of the seven, GPT-5.6 Sol leads one, and Grok leads two.

Grok 4.7 improves on Grok 4.6 in every row. That is the comparison the vendor controls completely, and the one that tells a buyer least.

Where it wins

The two wins are not marginal. On EEBench, an electrical engineering test, Grok 4.7 records 64.0 per cent against Fable's 56.4 and GPT-5.6 Sol's 39.4. On the Harvey legal agent benchmark it records 19.6 per cent against Fable's 6.7 and Sol's 2.5.

A threefold margin on a legal agent test could reflect genuine specialisation or a benchmark nobody else optimised for. The table cannot distinguish between the two, and neither can a reader.

2 of 7Benchmarks Grok 4.7 leads
4 of 7Led by Fable 5.1
37.6%Terminal-Bench, against Fable's 57.9
$2 / $6Input and output, per million tokens

Where it loses

Terminal-Bench 4.0 is the widest gap: 37.6 per cent against Fable 5.1's 57.9, a margin of twenty points on a test of working in a shell. On CursorBench it records 46.3 against Fable's 51.8, and on the multi-hour office-work test it scores 1,657 against 1,678.

HealthBench Professional puts it third: 56.7 per cent behind Sol's 60.5 and Fable's 62.1. On DeepSWE the three are within three points of each other, and Sol is ahead.

A model that trails on shell work and clinical reasoning while leading on electrical engineering and legal agents is not a general win. It is a shape.

Every figure is self-reported

xAI marks the results as its own testing. No independent verification is offered, and none of the comparison scores for rival models were produced by the companies that make them.

Reporting on the launch has been plain about this. OfficeChai, for instance, notes that the numbers come from the company itself.

Two of the figures also move depending on who reprints them. One aggregator lists EEBench at 66.0 per cent and Terminal-Bench at 38.0, against 64.0 and 37.6 on xAI's own page. The differences are small and the direction is consistent; that is a small lesson in how these numbers travel.

Advertisement

What xAI says changed

The company attributes the gains to a larger base model than 4.6, a longer reinforcement learning run weighted toward multi-hour problems, improved self-verification, better long-context handling, and training against its own agent harness.

Pricing is unchanged from the previous release at two dollars per million input tokens and six per million output, with a faster variant at double.

A vendor published a scorecard on which it comes third overall and first twice, and did not pretend otherwise. That is more disclosure than these launches usually carry; it is still a scorecard marked by the candidate.

Advertisement
Maya Lin
Digital Platforms Analyst

Maya Lin covers SaaS platforms, workflow automation, creator tools, and productivity software for RECATOOLS.

View author profile → · Editorial policy

About this byline Maya Lin is a RECATOOLS editorial persona used for platform and productivity coverage. Articles are produced and reviewed under RECATOOLS editorial supervision.

Corrections policy

Advertisement