
MODELS
by PrismML
A 27B-class ternary multimodal reasoning model from PrismML with 262K context, image input, and local GGUF and MLX releases.
Ternary Bonsai 2 27B is PrismML’s 27.8B-parameter multimodal reasoning model, derived from Qwen3.8-27B. PrismML describes it as a 5.9 GB ternary-compressed model with a 262K-token context window and text-and-image input. The model is positioned for reasoning, coding, vision, and agentic work. Official releases include a GGUF variant for llama.cpp on CUDA, Metal, and CPU, plus a 2-bit MLX repository. PrismML’s accompanying Bonsai Demo provides local deployment guidance and integrations such as configurable reasoning, image/PDF handling, and OpenAI-compatible tool-call endpoints.
Use for local coding mathematical reasoning and long-context tasks that can benefit from a 262K-token context window. Use the GGUF release when running llama.cpp on CUDA Metal or CPU hardware. Use the MLX 2-bit release for Apple Silicon or other MLX-based local workflows. Use the Bonsai Demo when building a local application with image or PDF inputs and OpenAI-compatible tool calling.
Last updated Sep 20, 2026