Prism ML releases Ternary-Bonsai-2-27B, 27B model compressed to 5.9GB
Prism ML has released Ternary-Bonsai-2-27B-gguf, a 27-billion parameter language model compressed to approximately 5.9GB using ternary quantization, down from 54GB in full precision format. The model retains 98.2% of full-precision performance across 14 thinking-mode benchmarks while running on consumer hardware, including Apple Silicon laptops at around 47 tokens per second on an M5 Max. The model uses a true 1.72 bits per weight through dense ternary packing and maintains the full 262K token context window from its Qwen base, with hybrid attention architecture enabling practical on-device inference. The release includes GGUF formats optimized for llama.cpp with CUDA and Metal support, plus separate variants for MLX and iOS deployment.
Alibaba releases Qwen3.8-27B, compact vision-language model with extended context
Alibaba has released Qwen3.8-27B, a 27-billion parameter vision-language model that builds on the Qwen3.5 architecture with improvements in coding, professional work, and agentic task completion. The model features a 262,144 token native context length extensible to 1 million tokens, native support for image and video understanding, and flexible reasoning controls including a thinking mode that can be tuned per request. Qwen3.8-27B introduces architectural innovations including Gated DeltaNet for linear attention alongside traditional gated attention mechanisms, and is designed as a compact, deployable option compared to larger models in the family. The model is available through Hugging Face Transformers and compatible with multiple inference frameworks including vLLM and SGLang, with a managed version coming soon through Qwen Cloud.