Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Yi-Large · Evaluations · Kaino
Yi-Large logo
model evaluation

Yi-Large

01.AI

01.AI’s hosted Yi-series flagship text model for chat/completions API use through an OpenAI-compatible developer platform.

modelllmyi-large
55.2KAINO SCORENot recommended
Evaluated Sep 15, 202610 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Technical capability75
  • Developer experience74
  • Risk & evidence65
  • Reasoning & knowledge64
  • Adoption signal55
  • Coding & agentic54
  • Cost effectiveness52
  • Speed & availability48
  • Multimodal & I/O40
  • Pricing clarity25

Kainotomic evaluation

Yi-Large has credible historical capability evidence, but the public record is dated and incomplete. The July 2024 Arena-Hard-Auto snapshot puts Yi-Large at 63.70, with Yi-Large-preview at 71.48; this supports a mid-tier general reasoning score rather than present-frontier positioning. Its reported HumanEval basic 0.524 and EvalPlus 0.652 indicate usable baseline coding ability, but do not establish repository-agent performance. It is materially below Claude Opus 4.8 and GPT-5.5 on evidenced frontier reasoning and coding, while its general profile is closer to Hermes 4.3 36B than GLM-4.6. Official sources substantiate hosted chat/completions access and OpenAI-compatible integration, supporting a workable developer-experience score. They do not substantiate image, audio, tool-use, structured-output, context-window, rate-limit, regional, latency, or reliability claims; multimodal/I/O and availability therefore score below text-API comparables such as GPT-5.6 Terra. The Yi GitHub repository is a positive public developer signal, but it is not a model-specific API reference for Yi-Large. Cost and pricing are deliberately conservative: the official material says flexible pricing but supplies no rates, billing units, or quotas. DeepSWE lists no Yi-Large result; no SWE-bench evidence was found; LiveCodeBench has no Yi row; and the supplied Artificial Analysis source provides no Yi-Large intelligence, price, or speed data. Arena-Hard is useful public preference evidence but historical, so it should not be read as a current competitive ranking.

Strengths

  • Officially documented hosted chat/completions access with OpenAI-compatible integration.
  • Historical Arena-Hard-Auto result provides independent general-preference evidence.
  • Reported HumanEval and EvalPlus results provide a limited coding baseline.

Caveats

  • Arena-Hard evidence is a July 2024 snapshot and may not represent the currently served endpoint.
  • HumanEval/EvalPlus is not evidence of SWE-bench, terminal, or coding-agent performance.
  • The supplied official documentation is provider-level rather than a detailed Yi-Large model card.