xAI releases Grok 4.7
xAI reported a larger base model than Grok 4.6, up to 2.1 trillion parameters against a reported 1.5 trillion, at unchanged pricing; independent ranking placed it sixth on Artificial Analysis's GDPval-AA index.
- Models & capabilities
- Notable
xAI released Grok 4.7, a coding and agentic model the company said uses a larger base model than its predecessor, Grok 4.6, alongside a longer reinforcement-learning run and training weighted towards harder, longer-running tasks. Secondary coverage reported the base model at roughly 2.1 trillion parameters, up from about 1.5 trillion for Grok 4.6, though xAI’s own post did not disclose parameter counts. The model keeps a 500,000-token context window, matching Grok 4.6, and takes text and image input.
Pricing was unchanged from Grok 4.6, at $2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens, doubling above that threshold. Elon Musk, xAI’s founder, had flagged the model as imminent in August and repeatedly pushed the release date back over the following weeks before it shipped on 21 September. The model is available through xAI’s API, Grok’s own coding tools and third-party harnesses including Cursor.
Independent evaluation gave a more measured picture than xAI’s own comparisons. On Artificial Analysis’s GDPval-AA leaderboard, which scores models by Elo rating on professional knowledge-work tasks, Grok 4.7 ranked sixth, behind five Anthropic entries — two effort settings each of Claude Opus 5.5 and Claude Fable 5.1, plus Claude Opus 5 — making it the highest-placed non-Anthropic model in the leaderboard’s top tier. As with prior Grok releases, the wider comparison figures xAI cited alongside the launch were the company’s own reported scores rather than independently verified results. xAI’s own table showed Grok 4.7 nearly doubling Grok 4.6 on the Terminal-Bench 4.0 agentic-coding test (37.6% against 20.3%) and leading on Harvey’s legal-agent benchmark (19.6%) and an electrical-engineering test, while Claude Fable 5.1 led four of the seven benchmarks shown, and it omitted the older reasoning benchmarks, such as GPQA and AIME, that earlier launches had featured.
The release continued xAI’s pattern of frequent, incremental point updates — Grok 4.5 in July, 4.6 in August, 4.7 in September — after a public roadmap that had promised the model within three to four weeks of Grok 4.6’s launch but slipped by roughly a month.