Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
xLAM · Evaluations · Kaino
X
model evaluation

xLAM

Salesforce AI Research

Large Action Model family focused on function calling, tool use, and agentic developer applications.

xLAMsalesforce
58.3KAINO SCORENot recommended
Evaluated Jul 31, 20268 reviews
Website Docs GitHub

Scorecard

PricingMultimodalCostDev expTechnicalSpeedCodingReasoningRiskAdoption
  • Technical capability76
  • Developer experience74
  • Coding & agentic70
  • Risk & evidence70
  • Reasoning & knowledge66
  • Adoption signal57
  • Cost effectiveness52
  • Speed & availability48
  • Multimodal & I/O45
  • Pricing clarity25

Kainotomic evaluation

xLAM has credible, specialized evidence for action-oriented systems: Salesforce’s repository and paper report BFCL, τ-bench, ToolBench, ToolQuery, WebShop, BOLAA, AgentLite, and MINT-BENCH evaluations. This supports a technical score below broad frontier models such as Claude Opus 4.8, but a coding-and-agentic score materially above Gemma 3’s 42 because the project is explicitly designed and evaluated for tool use. It is not evidence of general coding-agent leadership: xLAM is absent from the supplied DeepSWE and LiveCodeBench checks, and no SWE-bench result is supplied. The official materials and public repository provide usable implementation/project documentation, making developer experience comparable to but below more fully productized API platforms. The supplied evidence does not establish multimodal inputs, hosted inference, throughput, context limits, pricing, or commercial availability. Accordingly, multimodal, speed, cost-effectiveness, and especially pricing-clarity scores remain below DeepSeek-V3.2 and GPT-5.6 Luna anchors, whose catalog evaluations have clearer deployment or price evidence. Public signal is moderate: Salesforce AI Research authorship, a paper, and a maintained official repository are meaningful, but there is no supplied LMArena/Arena-Hard preference result or adoption metric. Evidence quality is reasonably strong for the narrow function-calling claim, but vendor-reported benchmark claims should not be equated with independently comparable frontier performance. xLAM is best treated as a focused open research/model-family option rather than a verified general-purpose production leader.

Strengths

  • Official Salesforce research paper and repository.
  • Specialized function-calling and tool-use benchmark coverage.
  • Relevant architecture and materials for agentic application evaluation.

Caveats

  • No supplied DeepSWE, SWE-bench, or LiveCodeBench result.
  • No established multimodal capability in supplied documentation.
  • Hosted availability, latency, pricing, and license terms are not established.