PrismML releases Ternary Bonsai 2 27B, shrinking Qwen3.8 27B to 5.93 GB with 98.2% average retention

PrismML releases Ternary Bonsai 2 27B, shrinking Qwen3.8 27B to 5.93 GB with 98.2% average retention

N
News Editor
2026-09-18 18:15:27
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B that cuts model size to 5.93 GB from 53.80 GB for the FP16 build while keeping an average 98.2% of performance across 20 benchmarks. The model supports both text and image input, offers a 262K-token context window, and is designed to run on consumer hardware including the RTX 5090. PrismML said the weights are open-sourced under the Apache 2.0 license. The architecture matches Qwen3.8 27B, with 27.36B total parameters split across a 24.35B language backbone, 2.54B embedding and language modeling head, and a 0.47B vision module. According to the reported evaluation, retention topped 99% on math and coding tasks, while performance dropped more noticeably on long-horizon agent workloads such as Terminal-Bench 2.1. Reported decoding speeds reached 142.5 tokens per second on an RTX 5090, 96.7 token/s on an RTX 4090, and 46.8 token/s on an Apple M5 Max laptop.

PrismML has released Ternary Bonsai 2 27B, according to Techub News. The model is a ternary-weight version of Qwen3.8 27B, reducing size to 5.93 GB from 53.80 GB for the FP16 version while retaining an average 98.2% performance across 20 benchmark tests.

The model supports both text and image input, has a 262K-token context window, and can run on consumer hardware such as the RTX 5090. PrismML released the weights under the Apache 2.0 license.

Architecture and quantization

The architecture remains aligned with Qwen3.8 27B. Total parameters stand at 27.36B, including a 24.35B language backbone, 2.54B for the embedding layer and language model head, and a 0.47B vision module.

The model uses ternary quantization, with most weights limited to -1, 0, and +1, combined with shared FP16 scaling factors to achieve a higher compression ratio.

Benchmark and speed results

Reported evaluation results showed retention above 99% on math and coding tasks. Performance fell more noticeably on long-horizon agent tasks, including Terminal-Bench 2.1.

In reported tests, decoding speed reached 142.5 tokens per second on an RTX 5090, 96.7 token/s on an RTX 4090, and 46.8 token/s on an Apple M5 Max laptop.

The item cited MarkTechPost.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.