Quick Facts

  • Grok 4.6 carries 1.5 trillion parameters, uses a 500,000-token context window, and starts at $2.00 per million input tokens.
  • The model scored 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol Max, and claimed the top spot on the Databricks leaderboard.
  • SpaceXAI plans to follow Grok 4.6 with Grok 4.7 three to four weeks later, then Grok 5 before the end of 2026.

SpaceXAI released Grok 4.6 on August 12, 2026, one month after Grok 4.5 shipped in July. The model is the first flagship release to carry the SpaceXAI name after SpaceX acquired xAI in an all-stock deal valued at $1.25 trillion that closed February 2, 2026.

Grok 4.6 does not increase parameter count over its predecessor. It uses the same 1.5 trillion-parameter V9 foundation and delivers performance gains through improved supervised fine-tuning and reinforcement learning. xAI used Grok 4.5 to regenerate training trajectories across reasoning, agent harnesses, and domains including STEM, software engineering, and knowledge work.

Agentic Design at the Center

The headline behavioral change for developers is self-verification. During long agentic runs, Grok 4.6 checks its own work before moving to the next step, reducing error accumulation across multi-step tasks. xAI trained the model on agentic reinforcement learning tasks covering general coding, kernel optimization, web development, and computer-aided design.

The model supports text and image input with text-only output. Developers get function calling, structured outputs, and a configurable reasoning_effort parameter now offering four levels: low, medium, high (default), and a new xhigh setting. The knowledge cutoff is February 1, 2026.

Benchmark Results

On xAI’s launch table, Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and tied with GPT-5.6 Sol Max. The model trails Fable 5 Max by one point at that level.

On DeepSWE, it jumped from 54.0 to 65.9, and on APEX-Agents from 47.1 to 57.5. The model also posted a 1,753 Elo score on GDPval-AA v2, up from 1,526 for Grok 4.5, and claimed the top ranking on the Databricks leaderboard. Terminal-Bench v3.0 reached 26%, nearly double Grok 4.5’s 15.7%, though it remains last among the four models listed in xAI’s comparison table.

DeepSWE v1.1 landed at 65.9%, behind GPT-5.6 Sol Max at 73%. Long-horizon task completion is emerging as a real-world differentiator alongside the model’s benchmark scores.

Pricing Structure

Pricing starts at $2.00 per million input tokens, $0.50 per million cached input tokens, and $6.00 per million output tokens for prompts under 200,000 tokens. At or above that threshold, rates double to $4.00, $1.00, and $12.00 per million tokens, and the higher rate applies to all tokens in that request.

Enterprise buyers should not apply the lower headline rates when estimating costs for long-context workloads. The pricing jump at 200K tokens represents a meaningful cost increase for teams running extended agentic sessions against the full 500K-token window.

What Comes Next

During SpaceX’s Q2 2026 earnings call, Elon Musk outlined a tight model roadmap. Grok 4.7, carrying 2.1 trillion parameters, is expected three to four weeks after Grok 4.6. Grok 5 is targeted before year-end.

The rapid release cadence follows SpaceX’s $60 billion acquisition of agentic coding platform Cursor in June and comes ahead of SpaceX’s planned $75 billion IPO. SpaceXAI is positioning Grok as a developer and enterprise platform, with the Cursor acquisition and agentic model design pointing toward that strategy.

Read more: SpaceXAI releases flagship Grok 4.6 model with advanced reasoning capabilities

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.