Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Perceptron Mk1 · Evaluations · Kaino
P
model evaluation

Perceptron Mk1

Perceptron AI

Vision-language model for image and video understanding, OCR, object detection, event clipping, and embodied reasoning via API.

multimodalclosed-source
69.4KAINO SCORENot recommended
Evaluated Jul 31, 20268 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness89
  • Multimodal & I/O80
  • Developer experience78
  • Pricing clarity76
  • Technical capability72
  • Speed & availability67
  • Reasoning & knowledge65
  • Risk & evidence64
  • Adoption signal58
  • Coding & agentic45

Kainotomic evaluation

Perceptron Mk1 has documented text, image, and video input; text output; OCR, detection, event-clipping, and embodied-reasoning use cases; a 32K context window; streaming; and OpenAI-compatible chat completions with vision controls. This makes its multimodal I/O stronger than the lower-scoring Gemma 3 anchor’s documented deployment profile, but below Gemini 3.5 Flash, whose 92 reflects substantially broader established multimodal evidence. No independent quality benchmark supports the vendor’s frontier-comparability claim, so technical and reasoning scores remain below Gemini, Kimi K2.5, and Claude Opus 4.8. The listed OpenRouter price of $0.15/$1.50 per million input/output tokens is exceptionally low against published anchors, supporting a high cost-effectiveness score. Pricing clarity is lower because the supplied official materials do not publish an authoritative rate card; the price, 20 tokens/s throughput, and 5.93-second latency come from OpenRouter. The API documentation and familiar OpenAI-compatible interface support solid developer experience, though the closed model has a smaller public integration and operational track record than Gemini or Claude. Coding and agentic scores are deliberately conservative: DeepSWE and LiveCodeBench checked results do not list Mk1, SWE-bench has no direct evidence, and the supplied repository and launch release do not establish Terminal-Bench or Aider performance. Public adoption evidence is similarly thin, with no LMArena or Arena-Hard listing. Documentation substantiates capabilities, not comparative accuracy, reliability, safety, or availability; this creates a materially weaker evidence base than the established anchors.

Strengths

  • Native documented image and video inputs with OCR, detection, and event-oriented workflows
  • OpenAI-compatible streaming chat-completions API with vision controls
  • Very low OpenRouter-listed token pricing

Caveats

  • No independent capability, reasoning, coding, or agent benchmark result in supplied evidence
  • Official materials do not provide a public authoritative price card
  • Closed-source model with limited public adoption and reliability evidence