SPAWNSY

xAI Ships Grok 4.6: Matches GPT-5.6 Sol Max, at a Fifth of the Price

Grok 4.6 scored the same as GPT-5.6 Sol Max on the Artificial Analysis leaderboard, but costs five times less. We look at where xAI's model wins, where it loses, and why price might matter more than the benchmark here.

AuthorFlaviSPAWNSY Editorial Desk
PublishedAugust 15, 2026
Read time4 min
SectionTech
Views2,809
Share
xAI Ships Grok 4.6: Matches GPT-5.6 Sol Max, at a Fifth of the Price

xAI shipped Grok 4.6 on August 12 and immediately put it head to head with GPT-5.6 Sol Max: the same score on the independent Artificial Analysis Intelligence Index, at a fifth of the price. Matching the score while cutting the bill by a factor of five isn't the kind of detail OpenAI can just shrug off.

Same score, different price

Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, overtaking the popular Chinese open-weight model Kimi K3 and tying GPT-5.6 Sol Max. Only Anthropic's Claude Opus 5 and Claude Fable 5 rank higher. Pricing starts at $2 per million input tokens and $6 per million output tokens, rising to $4 and $12 per million for prompts above 200,000 tokens. That's roughly a fifth of what OpenAI charges for comparable output from GPT-5.6 Sol.

The model launched already available in the xAI API, in Grok Build, in Cursor, and through OpenRouter, Vercel, and Cloudflare, exactly where developers actually decide which model to plug into their tools. Showing up in that many places on day one signals xAI is targeting the same customer segment as OpenAI and Anthropic, not just consumers using the Grok app.

A tie that isn't really a tie

The headline score hides differences that matter more in practice than one number on a chart. Sol clearly wins the harder engineering tasks: DeepSWE (73% versus 65.9%) and Terminal-Bench (34.6% versus 26%), the tests that measure actually working with code and a terminal, not just understanding an instruction. Grok wins the professional knowledge-work side instead: GDPVal-AA, AA-Briefcase, and above all Harvey LAB, where Grok's score (15.8%) is six times Sol's (2.5%). Those are tasks simulating legal work, document analysis, and reasoning over long, dense text.

That split suggests the two labs trained with different priorities: xAI leaned harder into long, complex agentic tasks and legal or business-style text work, while OpenAI still holds the edge where you actually need to write and test working code. For anyone picking a model for a specific job, that's more useful information than the aggregate score.

What "long-running agent" actually means here

xAI describes Grok 4.6 as built for tasks that stretch over time rather than single prompts with an instant answer. In practice that means scenarios where the model gets a goal, plans its own steps, calls tools, checks intermediate results, and corrects course along the way instead of waiting for a new instruction after every step. That's the same direction Claude Code, Codex, and Gemini CLI are all moving in: the model stops being a chat window and starts acting like an employee you can hand a multi-hour task to and come back for the finished result.

The high Harvey LAB score fits the same story. It's a benchmark simulating real legal work, parsing multi-page contracts, finding precedent, reasoning over long dense documents, exactly the kind of task where a model's stamina on long context matters more than any single flash of a clever answer. A sixfold lead over Sol on that one test suggests xAI deliberately invested in this specific niche instead of trying to be best at everything at once.

Why price matters more than the benchmark here

A tie on the leaderboard at a fifth of the price changes the math for anyone building a product on top of someone else's model, not just testing it in a browser. A company building a coding agent or a document-analysis tool pays per token, and at scale a fivefold gap in output pricing can decide which vendor wins the contract, regardless of who scores a point or two higher on any single test.

xAI is building Grok 4.6 with a clear focus on long-running agents and more ambitious interactive and visual work, extending what started with Grok 4.5. That's in line with where the whole industry is heading: models get judged less and less on how well they answer a single question, and more and more on how they hold up across multi-step tasks spread over hours of work.

For the end user, that boils down to something simple: if you're building at any real scale, choosing between Grok 4.6 and GPT-5.6 Sol Max stops being a question of "which one's smarter" and becomes "which specific task are you doing, and what are you willing to pay for it." At an identical overall score, that second question starts mattering more than the first.

Comments

Discussion

Join the conversation around this story.

0 entries

Join the discussion

Sign in to comment and reply to other readers.

Sign in

No comments yet

Start the discussion first.

Read next

All posts