Google’s preview Gemini 3.1 Pro reasoning model for complex multimodal, coding, agentic, and long-context tasks.
Gemini 3.1 Pro is a strong frontier preview with documented multimodal, long-context, coding, and tool-use positioning. Artificial Analysis reports an Intelligence Index of 46 and a 1M-token context window; Google’s model card reports evaluations across reasoning, multimodal, agentic, multilingual, and long-context work. It rates above Gemma 3 and GLM-4.6 on overall technical breadth, and above Gemini 3.5 Flash on coding/agentic evidence, while remaining below GPT-5.5 and Claude Opus 4.8 on demonstrated top-end coding capability. Coding evidence is mixed rather than uniformly leading. Terminal-Bench 2.1 records 65.6%±1.7% with Terminus 2 and 65.8%±1.7% with Gemini CLI, supporting a high agentic-work score, but DeepSWE reports only 12%±2% on 113 tasks. Google documents SWE-bench methodology and results, but the supplied evidence does not provide the numeric score. Arena places it eighth in Text Overall at 1486±4, a solid but not leadership-level public-preference signal; LiveCodeBench has no direct listing. At $2/M input and $12/M output tokens, it is materially less economical than DeepSeek-V3.2 and Gemma 3, though less costly than premium flagship anchors. Reported throughput is 122.4 tokens/s, but 22.05-second TTFT limits interactive speed. Google’s API documentation, model card, and cloud integration support a high developer-experience score, while preview status, uneven independent coding results, and reliance on several provider-reported evaluations warrant caution.