Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Qwen/Qwen3.6-27B-FP8 · Evaluations · Kaino
Qwen/Qwen3.6-27B-FP8 logo
model evaluation

Qwen/Qwen3.6-27B-FP8

Qwen

Qwen3.6-27B-FP8 is a 27.78B-parameter Qwen image-text-to-text model variant distributed with FP8 artifacts.

modellead-sourcehugging-face-popular-modelssource:github.comqwenqwen3.627bfp8image-text-to-texthugging-facemodelscopesafetensors
79.0KAINO SCORERecommended
Evaluated Aug 6, 202610 reviews
GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness88
  • Pricing clarity84
  • Multimodal & I/O84
  • Developer experience84
  • Speed & availability82
  • Technical capability80
  • Coding & agentic79
  • Risk & evidence74
  • Reasoning & knowledge68
  • Adoption signal67

Kainotomic evaluation

Qwen3.6-27B-FP8 has credible 27B-class capability evidence: official materials document image, text, and video input; text output; function calling; structured output; and a native 262k context window, with self-hosted extension to about 1M tokens. Its reported 77.2% SWE-bench Verified result supports a coding score above GPT-5.6 Luna’s 78 only narrowly, but it remains below Kimi K2.5 (86) and Claude Opus 4.8 (95), which have stronger broader agentic evidence. Artificial Analysis’ Intelligence Index of 37 does not justify a high reasoning score. Multimodal I/O is comparable to Kimi K2.5 and above Luna, given documented image and video inputs, but below Gemini 3.5 Flash’s 92 because the supplied evidence does not establish similarly broad modality quality. At a reported $0.61 blended per-million-token price through DeepInfra FP8, cost efficiency exceeds Luna’s 86-calibrated profile. Official Model Studio pricing, limits, tool support, plus vLLM/SGLang OpenAI-compatible examples make pricing and developer experience relatively strong. Reported throughput near 58–59 output tokens/s and 1.07-second TTFT support an above-median speed score. Evidence quality is moderate rather than leading: the key SWE-bench claim is vendor-reported, while third-party performance and pricing are provider-specific. No matching DeepSWE result or Arena coding-preference rank was found. LiveCodeBench, Terminal-Bench, and Aider were checked, but supplied sources provide no usable exact-model results; adoption therefore remains below established frontier API models.

Strengths

  • Officially documented multimodal inputs, 262k context, function calling, and structured output.
  • FP8 weights with vLLM and SGLang OpenAI-compatible serving paths.
  • Low reported blended inference price and strong reported serving throughput.
  • Vendor-reported 77.2% SWE-bench Verified result for the underlying 27B model.

Caveats

  • SWE-bench evidence is provider-reported and FP8 equivalence is stated rather than independently benchmarked.
  • Artificial Analysis measurements and price reflect a specific DeepInfra FP8 deployment, not every endpoint or self-hosted configuration.
  • The supplied evidence does not provide exact-model LiveCodeBench, Terminal-Bench, Aider, DeepSWE, or Arena results.