In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 layers never cache a token. 2.24x: MTPLX's multi-token prediction on an M5 Max, with the output distribution unchanged. 2.1 points: the MMLU cost of abliteration, published by exactly one of four builds. Every figure is attributed to whoever measured …
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.