Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Qwen/Qwen3-4B-Instruct-2507 · Evaluations · Kaino
Qwen/Qwen3-4B-Instruct-2507 logo
model evaluation

Qwen/Qwen3-4B-Instruct-2507

Qwen

Qwen3-4B-Instruct-2507 is a 4B-parameter multilingual instruct language model from Qwen for text generation.

modellead-sourcehugging-face-popular-modelssource:github.comtext-generationinstruct-modelmultilingualqwen34b-parametershugging-facegithub
64.5KAINO SCORENot recommended
Evaluated Aug 6, 20268 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness89
  • Developer experience80
  • Technical capability75
  • Risk & evidence70
  • Adoption signal67
  • Speed & availability62
  • Reasoning & knowledge57
  • Multimodal & I/O55
  • Coding & agentic48
  • Pricing clarity42

Kainotomic evaluation

Qwen3-4B-Instruct-2507 is a compact multilingual text instruction model with a reported 262k-token context window. Its Artificial Analysis Intelligence Index of 7 and 4B scale support a materially lower technical-capability assessment than Gemini 3.5 Flash, Kimi K2.5, GPT-5.5, and Claude Opus 4.8. It is text-oriented in the supplied material; there is no verified image, audio, video, or structured tool-I/O evidence, so multimodal scoring remains near the catalog floor rather than comparable multimodal anchors. Coding evidence is limited but concrete: the cited H4 report gives the base model a 35.1 LiveCodeBench result; its 40.3 figure applies to a separately Codeforces-SFT-trained variant and is not credited here. That places it above the minimally evidenced Pegasus coding profile but well below Kimi K2.5, GPT-5.6 Luna, and frontier coding anchors. No exact-model DeepSWE score is listed, and supplied SWE-bench, Terminal-Bench/Aider, and LMArena/Arena-Hard checks do not provide a usable result. Reasoning is therefore conservative. The official Qwen documentation, Hugging Face presence, and public GitHub repository make local integration and inspection comparatively straightforward. Artificial Analysis lists $0.00 per million input/output tokens, giving unusually strong potential cost effectiveness, but official provider pricing and an output-speed measurement were not supplied; this sharply limits pricing and availability confidence. Public release visibility is credible but weaker evidence of adoption than the major proprietary anchors.

Strengths

  • Verified 35.1 base-model LiveCodeBench result.
  • 262k-token context reported by Artificial Analysis.
  • Open documentation, Hugging Face distribution, and Qwen GitHub repository.
  • Reported zero token price creates strong potential cost efficiency.

Caveats

  • The 40.3 LiveCodeBench result is for a Codeforces-SFT variant, not the base model.
  • No supplied evidence verifies multimodal inputs, tool use, or agent orchestration.
  • No output-speed figure is available in the cited Artificial Analysis entry.
  • Official pricing and license terms were not provided.