xAI shipped Grok 4.6 on August 12 and immediately put it head to head with GPT-5.6 Sol Max: the same score on the independent Artificial Analysis Intelligence Index, at a fifth of the price. Matching the score while cutting the bill by a factor of five isn't the kind of detail OpenAI can just shrug off.
Same score, different price
Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, overtaking the popular Chinese open-weight model Kimi K3 and tying GPT-5.6 Sol Max. Only Anthropic's Claude Opus 5 and Claude Fable 5 rank higher. Pricing starts at $2 per million input tokens and $6 per million output tokens, rising to $4 and $12 per million for prompts above 200,000 tokens. That's roughly a fifth of what OpenAI charges for comparable output from GPT-5.6 Sol.
The model launched already available in the xAI API, in Grok Build, in Cursor, and through OpenRouter, Vercel, and Cloudflare, exactly where developers actually decide which model to plug into their tools. Showing up in that many places on day one signals xAI is targeting the same customer segment as OpenAI and Anthropic, not just consumers using the Grok app.
A tie that isn't really a tie
The headline score hides differences that matter more in practice than one number on a chart. Sol clearly wins the harder engineering tasks: DeepSWE (73% versus 65.9%) and Terminal-Bench (34.6% versus 26%), the tests that measure actually working with code and a terminal, not just understanding an instruction. Grok wins the professional knowledge-work side instead: GDPVal-AA, AA-Briefcase, and above all Harvey LAB, where Grok's score (15.8%) is six times Sol's (2.5%). Those are tasks simulating legal work, document analysis, and reasoning over long, dense text.
That split suggests the two labs trained with different priorities: xAI leaned harder into long, complex agentic tasks and legal or business-style text work, while OpenAI still holds the edge where you actually need to write and test working code. For anyone picking a model for a specific job, that's more useful information than the aggregate score.
What "long-running agent" actually means here
xAI describes Grok 4.6 as built for tasks that stretch over time rather than single prompts with an instant answer. In practice that means scenarios where the model gets a goal, plans its own steps, calls tools, checks intermediate results, and corrects course along the way instead of waiting for a new instruction after every step. That's the same direction Claude Code, Codex, and Gemini CLI are all moving in: the model stops being a chat window and starts acting like an employee you can hand a multi-hour task to and come back for the finished result.
The high Harvey LAB score fits the same story. It's a benchmark simulating real legal work, parsing multi-page contracts, finding precedent, reasoning over long dense documents, exactly the kind of task where a model's stamina on long context matters more than any single flash of a clever answer. A sixfold lead over Sol on that one test suggests xAI deliberately invested in this specific niche instead of trying to be best at everything at once.
Why price matters more than the benchmark here
A tie on the leaderboard at a fifth of the price changes the math for anyone building a product on top of someone else's model, not just testing it in a browser. A company building a coding agent or a document-analysis tool pays per token, and at scale a fivefold gap in output pricing can decide which vendor wins the contract, regardless of who scores a point or two higher on any single test.
xAI is building Grok 4.6 with a clear focus on long-running agents and more ambitious interactive and visual work, extending what started with Grok 4.5. That's in line with where the whole industry is heading: models get judged less and less on how well they answer a single question, and more and more on how they hold up across multi-step tasks spread over hours of work.
For the end user, that boils down to something simple: if you're building at any real scale, choosing between Grok 4.6 and GPT-5.6 Sol Max stops being a question of "which one's smarter" and becomes "which specific task are you doing, and what are you willing to pay for it." At an identical overall score, that second question starts mattering more than the first.





Comments
Discussion
Join the conversation around this story.
Join the discussion
Sign in to comment and reply to other readers.
Sign inNo comments yet
Start the discussion first.