Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Discover / Models

AI Models

Every model we track, with what it is for, who publishes it, and how it has scored in the evaluations we run. Open any entry for its capabilities, pricing signals and sources.

AI AgentsMCP ServersSkills

125 entries

  • Qwen: Qwen3.6 35B A3BQwenOpen-weight multimodal Qwen model with 35B total parameters and 3B active parameters per token.
  • Qwen: Qwen3.6 Max PreviewQwenHosted proprietary preview text-generation model from Qwen/Alibaba Cloud, available through Qwen Studio and the Alibaba Cloud Model Studio API as qwen3.6-max-preview.
  • DeepSeek: DeepSeek V4 FlashDeepSeekDeepSeek-V4-Flash is a DeepSeek API model with 1M-token context, 384K maximum output, tool calls, JSON output, and OpenAI ChatCompletions and Anthropic-compatible access.
  • Qwen: Qwen3.5 Plus 2026-04-20Qwen / AlibabaQwen3.5 Plus 2026-04-20 is a Qwen Cloud Qwen3.5-Plus snapshot from Alibaba with a 1M-token context window, multimodal input support, and text output.
  • Llama 4 Maverick 17B-128EMetaOpen-weight, natively multimodal mixture-of-experts Llama 4 model with 17B active parameters, 128 experts, and a 1M-token context window.
  • Qwen/Qwen3.6-27B-FP8QwenQwen3.6-27B-FP8 is a 27.78B-parameter Qwen image-text-to-text model variant distributed with FP8 artifacts.
  • deepseek-ai/DeepSeek-R1deepseek-aiDeepSeek-R1 is an open-source DeepSeek text-generation model with an official GitHub repository and availability through DeepSeek website and API access.
  • Qwen/Qwen2-1.5B-InstructQwenInstruction-tuned 1.5B-parameter Qwen2 text-generation model from Qwen.
  • Qwen/Qwen3-4B-Instruct-2507QwenQwen3-4B-Instruct-2507 is a 4B-parameter multilingual instruct language model from Qwen for text generation.
  • Qwen/Qwen3-8BQwenQwen3-8B is a Qwen text-generation model in the Qwen3 series, developed by the Qwen team at Alibaba Cloud.
  • BAAI/bge-base-en-v1.5BAAIEnglish BGE v1.5 base embedding model for feature extraction, search, and RAG retrieval workflows.
  • FLUX.2Black Forest LabsCurrent Black Forest Labs image model family for generation and editing, with pro, max, flex, and klein variants for production and local use.
  • openai-community/gpt2openai-communityGPT-2 checkpoint in the openai-community Hugging Face namespace, categorized for text generation.
  • Jamba Mini 2AI21 LabsEfficient 256K-context text model in AI21 Labs’ Jamba family for enterprise workflows, RAG, grounded QA, and deployable developer applications.
  • Anthropic: Claude Opus 4.6 (Fast)AnthropicFast mode for Claude Opus 4.6 is a faster inference configuration enabled with `speed: "fast"` for supported Opus models.
  • embed-v4.0CohereCohere’s multimodal embedding model for text, images, and mixed documents, with 128K context and configurable embedding dimensions for enterprise search/RAG.
  • Pegasus 1.5TwelveLabsTwelveLabs Pegasus 1.5 is a video-language model for prompt-based video analysis, segmentation, summaries, and time-based metadata through the TwelveLabs API.
  • Pika 2.2PikaPika 2.2 is a video generation model available through Pika’s official fal.ai API partnership for text-to-video and image-to-video workflows.
  • HiDream-O1-Image-Dev-2604HiDream AIMIT-licensed image generation foundation model for text-to-image, editing, and subject-driven personalization up to 2048px.
  • GPT-5OpenAIOpenAI model for coding and agentic tasks with text and image input, structured outputs, and parallel tool calling.
  • Phi-4-reasoning-vision-15BMicrosoftCompact 15B open-weight multimodal reasoning model for vision-language tasks, UI grounding, math, science, and document understanding.
  • Amazon Nova 2 LiteAmazonCost-efficient multimodal reasoning model on Amazon Bedrock for everyday automation, document processing, customer support, and agentic AI applications.
  • Deepgram Nova-3DeepgramHigh-accuracy Deepgram speech-to-text model for batch and streaming transcription, with multilingual support and self-serve terminology customization.
  • Adobe Firefly Image Model 5AdobeAdobe Firefly Image Model 5 is Adobe’s latest Firefly image generation model for API workflows, with emphasis on realism, lighting, composition, native 4 MP output, and natural-language instruct editing.
  • Morph V3 LargeMorphSpecialized code-apply model for merging AI-generated code edits with high accuracy, aimed at coding agents and developer-product workflows.
  • Sarvam-105BSarvam AIOpen-source 105B-parameter MoE reasoning chat model with an OpenAI-compatible API, long-context support, coding, agentic, and Indian-language strengths.
  • Krea 2 LargeKrea AIKrea 2 Large is Krea AI’s large variant of its in-house Krea 2 image foundation model, focused on aesthetics, style references, moodboards, and creative control.
  • Recraft V4.1 ProRecraftHigh-resolution, design-focused text-to-image generation model from Recraft, available through Recraft Studio and the Recraft API with raster and vector variants.
  • Bria 3.2BriaCommercial-safe text-to-image model trained on licensed data, with API support for generation and editing workflows.
  • Fish Audio S2 ProFish AudioExpressive text-to-speech model with natural-language emotion and prosody control, available through Fish Audio app/API and an open-source release.
  • SIMBA 3.0SpeechifyProduction voice AI model family for TTS, STT, and speech-to-speech via the Speechify Voice API, optimized for low latency and long-form stability.
  • EVI 4-miniHume AIMultilingual real-time speech-language model for emotionally intelligent voice interfaces with streaming speech and prosody-aware responses.
  • EXAONE 4.5LG AI ResearchOpen-weight 33B vision-language model from LG AI Research for text-image reasoning, document understanding, STEM tasks, and Korean contextual reasoning.
  • LongCat-NextMeituan LongCatOpen-source native discrete multimodal model unifying image, audio, and text in one autoregressive token space for multimodal generation and understanding.
  • Sonic 3.5CartesiaSonic 3.5 is Cartesia’s fast, natural streaming TTS model with sub-90ms latency, 42-language support, and stable dated snapshots.
  • Perceptron Mk1Perceptron AIVision-language model for image and video understanding, OCR, object detection, event clipping, and embodied reasoning via API.
  • Higgs Audio v3 TTSBoson AI4B conversational text-to-speech model for low-latency voice agents with multilingual speech, voice cloning, and inline emotion/style control.
  • Claude FableAnthropicClaude Fable 5 is an Anthropic Claude model described as a Mythos-level model for ambitious long-running work.
  • Arctic-ExtractSnowflakeSnowflake vision-language model for structured information extraction from documents, images, and text inside Cortex AI.
  • ZONOS2ZyphraOpen-source MoE text-to-speech model for real-time, high-fidelity voice cloning with Apache 2.0 weights and hosted or self-hosted inference paths.
  • Gemini 3.1 ProGoogleGoogle’s preview Gemini 3.1 Pro reasoning model for complex multimodal, coding, agentic, and long-context tasks.
  • Hermes 4.3 36BNous ResearchOpen-weight 36B hybrid-reasoning language model from Nous Research with long-context support and local-deployment-oriented availability.
  • Ring-2.6-1TInclusionAI / Ant GroupOpen trillion-parameter reasoning model for complex agentic workflows, coding, research, and enterprise automation.
  • DBRXDatabricksDBRX is an open general-purpose MoE language model from Databricks/MosaicML for coding, reasoning, and enterprise LLM use cases.
  • Kimi K2.7 CodeMoonshot AIOpen-weight, coding-focused Kimi model for long-horizon software engineering, tool use, large-context codebase work, and coding agents.
  • Qwen3-Coder-NextAlibaba QwenOpen-weight Qwen coding-agent model for software engineering, tool calling, structured outputs, and long-context developer workflows.
  • Llama 4 MaverickMetaMeta Llama 4 Maverick is an open-weight, natively multimodal Mixture-of-Experts model for text and image understanding.
  • MiniMax M2.5MiniMaxCoding- and agent-focused MiniMax text model with reasoning, tool use, high-throughput API variants, and open weights for local deployment.
  • xLAMSalesforce AI ResearchLarge Action Model family focused on function calling, tool use, and agentic developer applications.
  • gpt-oss-20bOpenAIOpenAI open-weight 20B model for local and developer workflows, including coding, reasoning, and tool-use experiments.
  • Gemini 3.5 FlashGoogle DeepMindFast Gemini 3.5 model for reasoning, coding, long-context, multimodal, and agentic tool-use workflows.
  • Step 3.5 FlashStepFunOpen sparse MoE reasoning model from StepFun with 196B total parameters, 11B active parameters, 256K context, and support for coding, tool use, agentic workflows, deep research, API access, and local deployment.
  • Kimi K3Moonshot AIMoonshot AI flagship 2.8T-parameter model with a 1M-token context window, native visual understanding, reasoning, long-horizon coding, and knowledge-work capabilities.
  • OpenHands LMAll Hands AI / OpenHandsOpenHands LM 32B v0.1 is an open coding-agent model from OpenHands / All Hands AI for autonomous software-development tasks.
  • MiMo-V2.5-ProXiaomiXiaomi’s flagship MiMo-V2.5-Pro large language model for complex agentic, reasoning, and coding tasks, available through Xiaomi MiMo and as open weights on Hugging Face.
  • Muse SparkMetaMuse Spark is a Meta Superintelligence Labs model family described by Meta as focused on multimodal reasoning, tool use, coding, and agentic tasks.
  • Amazon Nova PremierAmazon Web ServicesAWS frontier multimodal foundation model for complex reasoning, coding, tool use, RAG, and agentic enterprise workloads on Bedrock.
  • Granite 4.0 H SmallIBMIBM Granite-4.0-H-Small is a 32B total, 9B active hybrid MoE instruction model for enterprise RAG, agents, coding, tool use, and multilingual dialog.
  • Nemotron 3 UltraNVIDIANVIDIA Nemotron 3 Ultra is an open 550B-parameter MoE language model with 55B active parameters, 1M context, configurable reasoning, and tool-use support.
  • Grok 3xAIxAI's Grok 3 model family for coding, technical Q&A, and developer workflows via the xAI API.
  • Mistral Medium 3.5Mistral AIFrontier-class Mistral model optimized for agentic and coding use cases, with reasoning support and multimodal capabilities for developers.
  • GPT-5.4OpenAIOpenAI frontier model for professional reasoning, coding, tool use, and agentic workflows across ChatGPT, API, and Codex.
  • GLM-5.2Z.AIZ.AI flagship text foundation model for long-horizon coding and agentic engineering, with 1M context, 128K output, thinking modes, and MIT-licensed open weights.
  • GPT-5.6 LunaOpenAIOpenAI’s fastest, lowest-cost GPT-5.6 model for cost-sensitive, high-volume text, coding, reasoning, and tool-use workloads via the OpenAI API.
  • GPT-5.6 SolOpenAIOpenAI’s GPT-5.6 flagship model for advanced reasoning, coding, and complex agentic work.
  • GPT-5.6 TerraOpenAIGPT-5.6 Terra is described by OpenAI source metadata as a balanced GPT-5.6 model for everyday developer and enterprise workloads, positioned as a lower-cost option.
  • Yi-Lightning01.AI01.AI Yi-series flagship LLM available through the Lingyiwanwu developer platform.
  • Gemma 3Google DeepMindOpen model family from Google DeepMind for text, image understanding, reasoning, multilingual use, and developer deployment.
  • DeepSeek-V3.2DeepSeekReasoning-first open model built for agents, with thinking integrated into tool use and support for coding, reasoning, and instruction-following workflows.
  • LFM2.5-8B-A1BLiquid AILiquid AI compact MoE reasoning model for on-device tool calling, function calling, instruction following, and agentic tasks.
  • Yi-Large01.AI01.AI’s hosted Yi-series flagship text model for chat/completions API use through an OpenAI-compatible developer platform.
  • nomic-embed-text-v2-moeNomic AIOpen-source multilingual Mixture-of-Experts text embedding model for retrieval and RAG.
  • ERNIE 4.5BaiduBaidu’s ERNIE 4.5 model family includes text and vision-language variants for chat, tool use, structured output, and open deployment.
  • NVIDIA Alpamayo 2 SuperNVIDIA32B open reasoning VLA model in NVIDIA’s Alpamayo family for autonomous-vehicle research, trajectory generation, and chain-of-causation reasoning.
  • Ideogram 4.0IdeogramOpen-weight 9.3B text-to-image foundation model for design workflows, multilingual text rendering, layout control, and 2K image generation.
  • Olmo 3Ai2Fully open 7B/32B language model family with Base, Instruct, Think, and RL variants for chat, reasoning, tool use, and research workflows.
  • ElevenLabs Scribe v2ElevenLabsSpeech-to-text model for transcription across 90+ languages with word-level timestamps, speaker diarization, keyterm prompting, and audio tagging.
  • GLM-4.6Z.aiZ.ai open-weight LLM for coding, long-context reasoning, search, writing, and agentic development workflows.
  • Runway Gen-4.5RunwayRunway video generation model available in the Runway API as `gen4.5`, supporting text or image inputs for cinematic AI video generation.
  • Ray2Luma AILarge-scale generative video model for realistic visuals, coherent motion, and text/image-to-video workflows through Luma’s API.
  • Sonar Reasoning ProPerplexityWeb-grounded reasoning model for complex multi-step analysis and research tasks via Perplexity’s Sonar API.
  • LTX-2LightricksOpen audio-video foundation model for synchronized video and sound generation from Lightricks.
  • Stable Diffusion 3.5 LargeStability AIOpen text-to-image diffusion model focused on image quality, typography, complex prompt adherence, and local/custom deployment workflows.
  • Kimi K2.5Moonshot AIOpen-source multimodal agentic model for coding, visual understanding, documents, research, and multi-step developer workflows.
  • voyage-4-largeVoyage AIFlagship general-purpose and multilingual embedding model for high-quality retrieval, RAG, and search applications.
  • mistralai/Mistral-7B-Instruct-v0.3Mistral AIOpen 7B Mistral AI instruct model with 32k context, listed on the official Mistral 7B v0.3 model card and hosted on Hugging Face.
  • Qwen/Qwen3-1.7BQwenQwen3-1.7B is an open-weight dense Qwen3 text-generation model from the Qwen team at Alibaba Cloud.
  • sentence-transformers/paraphrase-multilingual-mpnet-base-v2sentence-transformersMultilingual Sentence Transformers model for semantic sentence similarity.
  • Qwen/Qwen3-VL-8B-InstructQwenQwen/Qwen3-VL-8B-Instruct is an image-text-to-text vision-language model from the Qwen3-VL family by the Qwen team at Alibaba Cloud.
  • Qwen/Qwen3-Embedding-0.6BQwenQwen3-Embedding-0.6B is a Qwen text embedding model listed for feature extraction and included in the Qwen3 Embedding model series.
  • meta-llama/Llama-3.1-8B-Instructmeta-llamaMeta’s Llama 3.1 8B Instruct is an instruction-tuned text-generation model in the Llama 3.1 family, listed on Hugging Face and documented by Meta’s official llama-models repository.
  • BAAI/bge-reranker-v2-m3BAAIMultilingual lightweight cross-encoder reranker from BAAI’s BGE reranker v2 family.
  • OpenRouter: FusionOpenRouterOpenRouter Fusion is a beta multi-model deliberation capability that runs multiple models in parallel and fuses their results into a single response.
  • hexgrad/Kokoro-82MhexgradKokoro-82M is an open-weight text-to-speech model from hexgrad with 82 million parameters.
  • laion/clap-htsat-fusedLAIONLAION CLAP model checkpoint for multimodal audio-and-language representation and audio-classification workflows.
  • Qwen/Qwen3-4BQwenQwen3-4B is a 4B-parameter dense text-generation model in the Qwen3 family from the Qwen team at Alibaba Cloud.
  • autogluon/chronos-2AutoGluonChronos-2 is a time-series forecasting model available on Hugging Face and documented for use with AutoGluon TimeSeriesPredictor.
  • pyannote/speaker-diarization-3.1pyannotepyannote/speaker-diarization-3.1 is a Hugging Face-hosted speaker diarization pipeline from pyannote, with usage documented via pyannote.audio and Pipeline.from_pretrained.
  • Qwen: Qwen3.7 PlusQwenQwen3.7-Plus is a Qwen multimodal agent model that supports text and image input with text output.
  • google/gemma-4-26B-A4B-itGoogleGoogle’s Gemma 4 26B A4B IT is a 26B-class Mixture-of-Experts Gemma 4 model variant for image-text-to-text workflows, with 25.2B total parameters, 3.8B active parameters, and a 256K context window.
  • ByteDance Seed: Seed-2.0-LiteByteDance SeedSeed2.0 Lite is a ByteDance Seed model in the Seed2.0 series, positioned to balance output quality and response speed for production-grade use.
  • Arcee AI: Trinity Large ThinkingArcee AIOpen reasoning model from Arcee AI for long-horizon agents, multi-turn tool calling, and reasoning tasks.
  • xAI: Grok 4.20 Multi-AgentxAIGrok 4.20 Multi-Agent Beta is an xAI API model for text/image-input, large-context, multi-agent research workflows.
  • Qwen: Qwen3.5-FlashQwen / Alibaba CloudQwen3.5-Flash is a Qwen flagship model with text, image, and video inputs, text output, and a 1,000,000-token context window.
  • LiquidAI: LFM2-24B-A2BLiquid AILiquid AI’s LFM2-24B-A2B is a 24B-parameter Mixture-of-Experts instruct text model in the LFM2 family, with about 2B active parameters, 32K context, and tool-calling support.
  • DeepSeek: DeepSeek V4 ProDeepSeekDeepSeek-V4-Pro is a DeepSeek Mixture-of-Experts language model with 1.6T total parameters, 49B activated parameters, and a 1M-token context window.
  • Anthropic: Claude Opus 4.7AnthropicClaude Opus 4.7 is Anthropic’s Opus-family model for long-running, asynchronous agents, with stronger coding and agentic performance than Opus 4.6.
  • Pareto Code RouterOpenRouterCoding-focused OpenRouter router that selects among strong coding models using Artificial Analysis coding percentiles.
  • MiniMax: MiniMax M3MiniMaxMiniMax M3 is a MiniMax frontier multimodal coding model with a 1M-token context window for agentic reasoning, tool use, coding, multimodal chat input, and long-context tasks.
  • Qwen: Qwen3.6 PlusQwen / Alibaba CloudQwen3.6-Plus is an API-accessible Qwen model with a listed 1M context window, 64k max output, and support for function calling, built-in tools, structured output, Coding Plan, reasoning, and visual understanding.
  • Google: Gemma 4 26B A4BGoogle DeepMindGemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts model from Google DeepMind with 25.2B total parameters and 3.8B active per token during inference.
  • Tencent: Hy3 previewTencentHy3 preview is Tencent’s open-sourced 295B-parameter MoE reasoning and agent model with 21B activated parameters and a 256K context window.
  • Xiaomi: MiMo-V2.5XiaomiNative omni-modal agent foundation model from Xiaomi MiMo with a 1M-token context window for image, video, audio, and text understanding.
  • Xiaomi: MiMo-V2.5-ProXiaomiMiMo-V2.5-Pro is Xiaomi’s MiMo V2.5 Pro Series language model for agent, coding, software-engineering, and long-horizon tasks, with a listed 1M-token context window.
  • OpenAI: GPT-5.5OpenAIGPT-5.5 is an OpenAI API model described for coding, professional work, complex production workflows, and tool-heavy use cases, with text/image input, text output, configurable reasoning effort, and token-based API pricing.
  • Qwen: Qwen3.6 FlashQwen / Alibaba CloudQwen3.6 Flash is a Qwen3.6 text-generation model available through Alibaba Cloud Model Studio as `qwen3.6-flash`, with a 1M-token context window.
  • Mistral: Mistral Medium 3.5Mistral AIMistral Medium 3.5 is Mistral AI’s open v26.04 frontier-class multimodal model, listed with a 256k context window and API pricing for input and output tokens.
  • inclusionAI: Ring-2.6-1TinclusionAIRing-2.6-1T is inclusionAI’s trillion-parameter flagship reasoning model for real-world complex task scenarios and agent workflows.
  • Perceptron: Perceptron Mk1Perceptron AIPerceptron Mk1 is a Perceptron AI vision-language model for image and video understanding with reasoning support.
  • xAI: Grok 4.3xAIGrok 4.3 is an xAI reasoning model with text and image inputs, text output, a 1,000,000-token context window, configurable reasoning, function calling, and structured outputs.
  • NVIDIA: Nemotron 3 Nano Omni (free)NVIDIANVIDIA Nemotron 3 Nano Omni is an open multimodal model for video, audio, image, and text workflows.
  • IBM: Granite 4.1 8BIBMIBM Granite 4.1 8B is an 8B-parameter long-context instruct language model in IBM’s Granite 4.1 family for general-purpose enterprise applications.
  • StepFun: Step 3.7 FlashStepFunHigh-efficiency multimodal sparse MoE vision-language model from StepFun, available through StepFun API and chat interfaces.
  • Google: Gemini 3.5 FlashGoogleGoogle Gemini 3.5 Flash is a high-efficiency multimodal Gemini API model with Flash-tier speed and cost positioning.
  • Anthropic: Claude Opus 4.8AnthropicClaude Opus 4.8 is Anthropic’s generally available Opus-family model for complex reasoning and agentic coding, with Claude API model ID `claude-opus-4-8`.