Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
EXAONE 4.5 · Evaluations · Kaino
E
model evaluation

EXAONE 4.5

LG AI Research

Open-weight 33B vision-language model from LG AI Research for text-image reasoning, document understanding, STEM tasks, and Korean contextual reasoning.

modelsource:lgresearch.aivision-language-modelopen-weight33bLG AI ResearchEXAONE
72.9KAINO SCORERecommended
Evaluated Jul 31, 20269 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Cost effectiveness87
  • Multimodal & I/O86
  • Technical capability81
  • Developer experience78
  • Reasoning & knowledge77
  • Coding & agentic75
  • Risk & evidence70
  • Adoption signal63
  • Speed & availability62
  • Pricing clarity50

Kainotomic evaluation

EXAONE 4.5 is a credible 33B open-weight VLM with a differentiated emphasis on document understanding, text-image reasoning, STEM, and Korean visual context. Its official technical report claims an 81.4 LiveCodeBench v6 result for the 33B Reasoning variant, but the current independent LiveCodeBench leaderboard does not list 4.5; this supports a solid, not top-tier, coding score. The official material substantiates multimodal availability, while public evidence is insufficient to place it near Claude Opus 4.8 or GPT-5.5 on broad capability. Relative to open-weight GLM-4.6, EXAONE scores higher in multimodal I/O because it is explicitly a released VLM with document-oriented positioning; its technical and reasoning scores are also modestly higher given the reported LiveCodeBench result. It remains below Gemini 3.5 Flash on multimodal ecosystem maturity, serving availability, and developer integration. The open weights and Artificial Analysis’ $0 estimated token pricing support strong cost effectiveness for operators able to self-host, but that estimate is not a provider API price. Evidence quality is moderate rather than high: central performance evidence is vendor-reported, and no DeepSWE/DataCurve, SWE-bench, Terminal-Bench, Aider, LMArena, or Arena-Hard result was supplied. LG’s official page and GitHub repository improve reproducibility and developer confidence, but supplied sources do not specify license terms, hosted-service pricing, measured inference speed, or broad adoption. Scores therefore favor the documented VLM strengths while discounting missing independent agentic and preference evidence.

Strengths

  • Open-weight 33B VLM with documented text-image and document-understanding focus.
  • Officially reported 81.4 LiveCodeBench v6 result for the 33B Reasoning variant.
  • Potentially strong self-hosted economics and a 262k-token context window listed by Artificial Analysis.
  • Korean contextual and visual-language specialization is a meaningful differentiator.

Caveats

  • The 81.4 LiveCodeBench score is vendor-reported; EXAONE 4.5 is absent from the checked official leaderboard.
  • No supplied independent software-engineering agent benchmark result supports deployment for repository-scale coding.
  • Artificial Analysis lists estimated zero token cost, not a confirmed LG hosted API tariff.
  • Speed, production availability, license terms, and serving requirements are not established by the supplied evidence.