Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
xAI’s Grok 4.7 Release Claims a 111-Elo Gain in Long-Horizon Agent Tests · News · Kaino
xAI’s Grok 4.7 Release Claims a 111-Elo Gain in Long-Horizon Agent Tests
Kaino
YesterdaySep 22, 2026, 12:00 AM33 views

xAI’s Grok 4.7 Release Claims a 111-Elo Gain in Long-Horizon Agent Tests

xAI says Grok 4.7 keeps its API pricing while Artificial Analysis records gains in two agent benchmarks.

xAIGrok 4.7Grok BuildAI agents

Grok 4.7 Just Changed the AI Coding Race — xAI’s New AI Model Brings 500K Context and a 111 Elo Jump xAI has released Grok 4.7, its latest AI model designed for extended coding, AI agent and knowledge-work tasks — and the benchmark numbers are already turning heads. According to Artificial Analysis, Grok 4.7 scored 111 Elo points higher than Grok 4.6 on AA-Briefcase, a benchmark focused on long-horizon agent work. The model also recorded a 9-point improvement on the Coding Agent Index when used with Grok Build. That makes Grok 4.7 more than a routine model update. It signals xAI’s growing focus on AI coding agents, long-running tasks and software development workflows. Grok 4.7 Has a Massive 500K Context Window One of the biggest features of Grok 4.7 is its 500,000-token context window. The xAI API documentation lists support for a 500K context window along with text and image inputs and configurable reasoning effort. For developers, a larger context window can be useful when working with large codebases, lengthy documentation, project history or complex multi-step tasks. But there is an important distinction: a 500K context window does not automatically mean perfect understanding of 500K tokens. The capability tells developers how much information the model can potentially process in context. It does not, by itself, prove that the model will consistently find the right information, complete every multi-step task correctly or avoid failures during tool use. Grok 4.7 Shows a 111 Elo Improvement The biggest benchmark headline is the 111 Elo gain over Grok 4.6 on Artificial Analysis' AA-Briefcase benchmark. AA-Briefcase evaluates models on longer-horizon agent tasks, making the result particularly relevant to the growing category of AI agents that can work through multi-step problems instead of simply generating individual responses. Artificial Analysis also reports a 9-point improvement on its Coding Agent Index when Grok 4.7 is used with Grok Build. These results suggest measurable improvement over Grok 4.6 in the tested environments. However, benchmark results should not be confused with a universal ranking of AI models. Performance can change depending on the coding framework, tools, prompts, repository, task complexity and agent setup. Grok 4.7 Is Targeting AI Coding Agents The direction of Grok 4.7 is particularly interesting for software developers. Traditional AI coding assistants typically help with individual tasks such as generating functions, explaining code or fixing errors. AI coding agents are increasingly designed to handle longer workflows — understanding a repository, making multiple changes, using tools, running tests and iterating on the result. Grok 4.7’s emphasis on extended coding and knowledge work puts it directly into this evolving category. The model is listed as available through the Grok API, Grok Build and Cursor, giving developers multiple ways to integrate it into their workflows. Grok 4.7 API Pricing Stays at $2/$6 There is another number developers will be watching closely. xAI's listed API pricing remains: $2 per million input tokens $6 per million output tokens That means developers already evaluating the xAI API can compare Grok 4.7's reported performance improvements against the same published token pricing. But API pricing isn't the same thing as total AI agent cost. A real-world AI coding agent can consume tokens through large context windows, repeated reasoning, tool calls, retries and generated code. Infrastructure and surrounding tools can also add to the total cost. So the real question for developers isn't simply how much Grok 4.7 costs per million tokens. It's how much it costs to successfully complete a task. Grok 4.7 vs Grok 4.6: What's Actually Different? The most notable changes reported so far are: 111 Elo improvement on AA-Briefcase 9-point Coding Agent Index improvement with Grok Build 500K-token context window Text and image input support Configurable reasoning effort Availability through Grok API, Grok Build and Cursor Listed API pricing of $2/M input and $6/M output Together, these features show where xAI is pushing Grok: toward long-context AI, autonomous coding workflows and agent-based software development. The Bigger AI Coding Story The most interesting part of Grok 4.7 may not be the model number itself. It's the direction of the AI industry. AI models are moving from "help me write code" toward "help me complete the task." That shift changes what developers expect from AI coding tools. Instead of asking an AI model to generate one function, developers can increasingly expect an agent to understand a larger project, reason through a problem, interact with tools and work through multiple steps. Grok 4.7 is another example of that transition. Its benchmark improvements provide evidence of progress, but they don't establish that it will outperform every competing model or every AI coding workflow. The more important test will happen in real repositories, real development teams and real production environments. The AI coding race isn't slowing down. It's moving from better answers to longer, more capable work.

Key takeaways
  • 1

    According to Artificial Analysis, Grok 4.7 scored 111 Elo points higher than Grok 4.6 on AA Briefcase, a benchmark focused on long horizon agent work.

  • 2

    The model also recorded a 9 point improvement on the Coding Agent Index when used with Grok Build.

  • 3

    It signals xAI’s growing focus on AI coding agents, long running tasks and software development workflows.

Continue reading

Latest from Kaino News

Story pulse

Freshness

Yesterday

Views

33

Reading

4 min

Byline

Kainotomic Team

Utilities

Topics

xAIGrok 4.7Grok BuildAI agents

Sources

Reference material and original reporting used in this story.

xAI

Published Sep 22, 2026, 12:00 AM

View source