Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Gemini 3.5 Flash · Evaluations · Kaino
Gemini 3.5 Flash logo
model evaluation

Gemini 3.5 Flash

Google DeepMind

Fast Gemini 3.5 model for reasoning, coding, long-context, multimodal, and agentic tool-use workflows.

google-deepmindlgemini-3.5-flash
82.4KAINO SCORERecommended
Evaluated Jul 31, 202610 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Multimodal & I/O92
  • Pricing clarity88
  • Developer experience88
  • Technical capability86
  • Speed & availability82
  • Reasoning & knowledge79
  • Coding & agentic78
  • Adoption signal78
  • Risk & evidence77
  • Cost effectiveness76

Kainotomic evaluation

Gemini 3.5 Flash has unusually broad documented API coverage: 1M-token input context, 65,536-token output, text/image/video/audio/PDF input, and function calling, structured outputs, grounding, URL context, code execution, and computer-use preview. This supports a multimodal/I/O score above Kimi K2.5 and GPT-5.5, and matches the previously published Gemini 3.5 Flash anchor. Artificial Analysis reports a 50 Intelligence Index, 169.6 output tokens/s, and the listed $1.50/$9.00 per-million input/output-token rates, supporting solid but not exceptional capability, speed, and value scores. Coding evidence is mixed. Google reports 55.1% on SWE-Bench Pro Public and DeepSWE lists 37% ±2%, rank 15 of 18 displayed entries; these results do not support placement near Claude Opus 4.8 or GPT-5.5 on agentic software work. A Kaggle LiveCodeBench listing reports 93.0% ±1.6% and rank one, but the current official LiveCodeBench page does not list this model, so that result receives limited weight. The resulting coding score is above the prior 71-point Gemini anchor due to additional benchmark evidence, but below Kimi K2.5 and DeepSeek-V3.2. Arena places the high variant eighth at 1480±6, a useful preference signal rather than proof of reasoning quality. Google’s mature Gemini API documentation, structured-output/tool support, and published standard pricing justify strong developer-experience and pricing-clarity scores. Evidence quality is reduced by benchmark-version inconsistency, absent Terminal-Bench/Aider results, and reliance on provider-reported SWE-Bench Pro performance.

Strengths

  • Documented 1M-token context and broad multimodal input support.
  • Strong Gemini API tool-use surface, including function calling, grounding, code execution, and structured outputs.
  • Published standard pricing and independently reported throughput data.

Caveats

  • DeepSWE performance is modest relative to leading coding-agent models.
  • LiveCodeBench evidence conflicts with the current official leaderboard listing.
  • Computer use is documented as preview functionality.