What is benchmaxing ?
Google recently announced their new model, Gemini 4 Argon, which performs incredibly well in benchmarks. That being said, how good they are in real life still remains a question we need to ask.
I've seen Gemini models get stuck thinking. Like, literally, the text in the thinking stream starts describing how it's stuck, and then it carries on. It feels dystopian watching that. And these are models that were getting great benchmark scores.
Sometimes I've seen Gemini fail in tool calls as well. So when another Gemini launch comes with good benchmarks, I don't immediately believe that the day-to-day problems are fixed. That's the problem I have with the excitement around Gemini 4 Argon. There are impressive numbers, but I still want to know whether using it is going to be frustrating.
Benchmaxing means chasing benchmark scores until the model becomes better at the tests without getting much better at the work people use it for. Say a team keeps choosing its next model based on which version does best on the same public test. That test starts shaping what gets built. Eventually you need to give the model unfamiliar work to find out whether it got better at the job or just better at those questions.
There's a more direct problem when test questions get into training data. Even changing the wording doesn't necessarily prevent it. A 2023 study showed that training on paraphrased and translated test examples could inflate benchmark results while evading string-matching checks for contamination. That's one way scores can become misleading; it isn't evidence that Google trained Argon this way.
Google obviously has good numbers to show. Argon scores 77.9% on DeepSWE v1.1 in its launch table. On FrontierSWE v2, though, it gets 55%, behind all three competitors in the comparison.
The Arena AI comparison I watched didn't give me confidence that Argon would be much better in use. The channel's own description says the generations were inconsistent across prompts.
And this gradual rollout feels like more hype. As of October 2, Google is starting with trusted cyber defenders and testers. Broader access is supposed to follow, beginning with paid API customers and Google AI Ultra subscribers, but the announcement doesn't give a firm date. We'll believe it when we see it, when we use it. Only then can we check whether it's still failing on the things we need it to do.
The one-million-token output limit is real. Google says it has raised the limit from 64K to give Argon more room for deep thinking and long tasks. That's impressive, but I don't know what I'd do with that much output in most of my day-to-day work. The problem I've had with Gemini is getting stuck thinking in the first place. More room to think doesn't tell me whether that's been fixed.
There is good research happening with these models too. Google's Gemini 2.0-based AI co-scientist proposed drug-repurposing candidates for acute myeloid leukemia that researchers validated in cell-line experiments. That's good stuff. It still doesn't tell me much about the coding assistant I'd be paying to use.
Google also says it's using Argon internally for code migrations and memory optimization. I'd be more interested in trying that sort of work myself than hearing more about the output limit. Developers buying access should be able to find out whether those results hold up for them. Until then, I think Google should hype it less and show more of the work getting done.
- 1
Like, literally, the text in the thinking stream starts describing how it's stuck, and then it carries on.
- 2
And these are models that were getting great benchmark scores.
- 3
So when another Gemini launch comes with good benchmarks, I don't immediately believe that the day to day problems are fixed.
Continue reading