Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
PrismML: Ternary Bonsai 2 27B · Discover · Kaino
Discover/MODELS/PrismML: Ternary Bonsai 2 27B
PrismML: Ternary Bonsai 2 27B logo

MODELS

PrismML: Ternary Bonsai 2 27B

by PrismML

modellead-sourcePrismML
Visit WebsiteDocumentationGitHub

Overview

A 27B-class ternary multimodal reasoning model from PrismML with 262K context, image input, and local GGUF and MLX releases.

Details

Ternary Bonsai 2 27B is PrismML’s 27.8B-parameter multimodal reasoning model, derived from Qwen3.8-27B. PrismML describes it as a 5.9 GB ternary-compressed model with a 262K-token context window and text-and-image input. The model is positioned for reasoning, coding, vision, and agentic work. Official releases include a GGUF variant for llama.cpp on CUDA, Metal, and CPU, plus a 2-bit MLX repository. PrismML’s accompanying Bonsai Demo provides local deployment guidance and integrations such as configurable reasoning, image/PDF handling, and OpenAI-compatible tool-call endpoints.

When to Use

Use for local coding mathematical reasoning and long-context tasks that can benefit from a 262K-token context window. Use the GGUF release when running llama.cpp on CUDA Metal or CPU hardware. Use the MLX 2-bit release for Apple Silicon or other MLX-based local workflows. Use the Bonsai Demo when building a local application with image or PDF inputs and OpenAI-compatible tool calling.

Getting Started

  1. Choose the GGUF release for llama.cpp or the MLX 2-bit release for MLX-based inference.
  2. Download the selected model files from PrismML’s official Hugging Face repository.
  3. Follow the setup guidance in the PrismML-Eng/Bonsai-demo repository for local deployment.
  4. For tool-enabled applications
  5. use the demo server’s documented OpenAI-compatible `/v1/chat/completions` endpoint.

Key Features

  • •27.8B-parameter ternary-compressed multimodal model
  • •262K-token context window
  • •Text and image input
  • •Approximately 5.9 GB model footprint
  • •GGUF release for llama.cpp on CUDA
  • •Metal
  • •and CPU
  • •2-bit MLX release
  • •Local Bonsai Demo with configurable reasoning and OpenAI-compatible tool-call support

Capabilities

  • •reasoning
  • •coding
  • •mathematics
  • •vision
  • •image understanding
  • •long-context inference
  • •agentic workflows
  • •local inference

Last updated Sep 20, 2026