SpaceXAI's New "Killer App": Grok 4.6 Agent Outperforms GPT-5.6 and Fable 5 in Tests, Priced at Half That of Competitors
I'm LongbridgeAI, I can summarize articles.The API for Grok 4.6 starts at $2 per million input tokens and $6 per million output tokens. The model enhances long-horizon execution capabilities, enabling autonomous testing and continuous correction. It has entered the top tier of frontier models in multiple agent and knowledge work benchmarks, achieving a score of 1753 Elo in the GDPVal-AA v2 test, surpassing GPT-5.6 Sol Max and Claude Fable 5 Max, and significantly outperforming the previous generation, Grok 4.5 High. SpaceXAI stated that during the first week of launch, users of Cursor and Grok Build will receive double the included usage allowance
On Wednesday, August 12 (Eastern Time), Elon Musk’s SpaceXAI officially released its next-generation large language model, Grok 4.6. SpaceXAI stated that the new model represents a significant improvement over Grok 4.5, with a key focus on enhancing capabilities in long-running agents, complex programming, knowledge work, and interactive and visual tasks.
The most notable aspect of this release is not merely the model iteration, but the fact that Grok 4.6 has joined the top tier in multiple frontier model tests published by SpaceXAI, while its API starting price is only $2 per million input tokens and $6 per million output tokens. SpaceXAI claimed its pricing is approximately half that of other frontier models; during the first week of launch, users of Cursor and Grok Build will also receive double the included usage allowance.
According to test data released by SpaceXAI, Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, tying with OpenAI’s frontier model, GPT-5.6 Sol Max. In the GDPVal-AA v2 benchmark, it achieved 1753 Elo, exceeding GPT-5.6 Sol Max’s 1728 and Anthropic’s frontier model, Claude Fable 5’s, 1741.

Elon Musk subsequently reposted social media posts stating that Grok 4.6 was a massive upgrade, scoring higher than Grok 4.5 High across all benchmarks. He emphasized Grok 4.6’s score of 1753 Elo and boasted that the new model was impressive.

From Knowledge Work to Programming, Grok 4.6 Comprehensively Enhances Complex Task Capabilities
The core of Grok 4.6’s upgrade is no longer just about answering questions, but rather enabling the model to continuously complete complex tasks requiring multi-step execution.
SpaceXAI stated that Grok 4.6 was specifically trained for long-running agents, capable of consecutively completing tasks such as researching topics, analyzing information, processing codebases, and transforming vague product ideas into runnable applications or deliverables. The company also noted that the new model begins to demonstrate greater autonomy in testing and validation over longer task trajectories, checking previous work before proceeding.
In terms of training, Grok 4.6 underwent more extensive supplementary training than Grok 4.5, incorporating filtered model-generated reasoning data and high-quality engineering data, along with improved optimizers and training schemes. Subsequently, SpaceXAI used Grok 4.5 to generate Supervised Fine-Tuning (SFT) trajectories covering various reasoning intensities, agent frameworks, and domains including STEM, software engineering, and knowledge work. The model was further refined through reinforcement learning to complete tasks such as general programming, knowledge work, kernel optimization, web development, and computer-aided design.
SpaceXAI particularly emphasized that Grok 4.6 has improved in turning “ideas” into “products.” For a specific product concept, the model can first research unfamiliar fields, design application structures, and implement core interactions, followed by multiple rounds of iteration based on feedback. In visual and interactive projects, the new model can also faster produce a relatively complete first version.
Low-Price Strategy Targets Developer Market, Doubling Allowances for Products Like Cursor in First Week
Compared to the model’s capabilities themselves, Grok 4.6’s pricing strategy may be of greater interest to the AI industry chain.
SpaceXAI announced API prices starting at $2 per million input tokens and $6 per million output tokens, with a faster, double-priced version also available. The company simultaneously announced that Grok 4.6 has been integrated into Cursor and Grok Build, and is available through partners such as OpenRouter, Vercel, and Cloudflare.
For comparison, SpaceXAI positioned Grok 4.6 in its announcement as “half the price of other frontier models.” Gavin Baker, Chief Investment Officer at Atreides Management, offered a more aggressive assessment: he believes Grok 4.6’s performance is roughly equivalent to Fable 5 Max, but with input token prices 80% lower and output token prices 88% lower, giving it a clear advantage in the combination of performance and cost.
Kim Monnis, a technology analyst and editor of Superintelligence Newsletter, also pointed out that Grok 4.6 is not only cheaper than Opus 5 and GPT-5.6 Sol, but even lower than Sonnet 5, while maintaining performance at least at the level of top-tier models. In his view, this combination of “frontier performance + low cost” is Grok’s true moat.
It should be noted that API unit prices do not fully equate to actual usage costs. Differences in token consumption, reasoning intensity, and task completion paths exist among different models. Therefore, “half the price” primarily refers to a comparison of published API prices, not implying that the final cost for all specific tasks is exactly 50% lower.
First-week promotions further reinforce this low-price strategy: SpaceXAI stated that Cursor and Grok Build users will receive double the included usage allowance during the first week, allowing developers to trial Grok 4.6 at a low cost.
Enhancing Long-Horizon Execution, Model Can Autonomously Test and Continuously Correct
From the test results, Grok 4.6’s improvements are particularly concentrated in the area of agents executing real-world work, which is becoming the core of competition among AI models.
SpaceXAI highlighted the GDPVal-AA v2 benchmark launched by Artificial Analysis. This test is based on OpenAI’s GDPval dataset and evaluates model performance in real-world, economically valuable professional tasks, covering 44 occupations and 9 major industries. Models are required to use tools such as Shell and web browsing within an agent environment to complete tasks, and final output files are compared via blind pairwise testing to generate an Elo score.
In the comparative data released by SpaceXAI, Grok 4.6 achieved 1753 Elo in GDPVal-AA v2, ranking higher than GPT-5.6 Sol Max’s 1728 and Claude Fable 5 Max’s 1741, a significant increase from the previous generation Grok 4.5 High’s 1526.
Meanwhile, Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, tying with GPT-5.6 Sol Max. This index is composed of nine tests and is used to measure a model’s comprehensive capabilities across dimensions such as mathematics, science, programming, and reasoning.
In other tests published by SpaceXAI, Grok 4.6 also showed significant progress compared to the previous generation:
- CursorBench v3.2: 69.9%, compared to Grok 4.5 High’s 66.7%;
- FrontierCode v1.1: 61.3%, compared to Grok 4.5 High’s 56.6%;
- APEX-Agents: 57.5%, compared to Grok 4.5 High’s 47.1%;
- APEX-SWE: 56.4%, compared to Grok 4.5 High’s 53.6%;
- AA-Briefcase: 1577 points, compared to Grok 4.5 High’s 1313;
- Harvey LAB: 15.8%, compared to Grok 4.5 High’s 12.9%.
However, Grok 4.6 does not rank first in all tests. For example, in DeepSWE 1.1 and Terminal-Bench v3.0, the scores for GPT-5.6 Sol Max published by SpaceXAI remain higher than those of Grok 4.6. Therefore, a more accurate statement is that Grok 4.6 has entered the top tier of frontier models in multiple agent and knowledge work tests, rather than comprehensively defeating all competitors.
It is worth noting that the Artificial Analysis GDPVal-AA v2 leaderboard itself changes with the addition of new models and updates to evaluations. Therefore, the term “outperforms” used in the title of this article specifically refers to the cross-sectional test results published by SpaceXAI at the time of this release, and does not claim that Grok 4.6 ranks first in all evaluations at all times.
From Model to Ecosystem, SpaceXAI Accelerates Capture of AI Development Entry Points
The release of Grok 4.6 also shows that SpaceXAI is extending competition from pure model performance comparisons to developer entry points and application ecosystems.
Currently, Grok 4.6 has simultaneously entered Cursor, Grok Build, and the API market, and is accessible through channels such as OpenRouter, Vercel, and Cloudflare. SpaceXAI hopes to leverage these platforms to allow developers to directly embed Grok 4.6 into programming, web development, and agent workflows, rather than just using the model through the Grok chat product.
Elon Musk himself has been actively promoting the new model on X. He first announced “Grok 4.6 is now out,” then reposted reports regarding the 1753 Elo ranking, emphasizing that Grok 4.6 took first place in GDPVal-AA v2. In another repost, he listed Grok 4.6’s rankings in tests such as AA-Briefcase, Harvey LAB, CursorBench, FrontierCode, APEX-Agents, and APEX-SWE, finally commenting: “Grok 4.6 is a banger.”
Gavin Baker of Atreides Management further directed market attention to the next generation of products. He predicted that Grok 4.7 will be a significantly larger-scale model and may incorporate Cursor and SpaceX data into pre-training. This statement currently represents the judgment of investors and is not a product roadmap confirmed in SpaceXAI’s release announcement.
If Grok 4.6 can maintain frontier-level capabilities while sustaining lower token prices, its significance may extend beyond the performance upgrade of a single generation of models. It could become a new signal that the price war in large AI models is extending into the agent era: as the capability gaps between models from different vendors gradually narrow, the balance between performance and inference costs may increasingly become the key factor for developers and enterprises in selecting models.
