Seen the same thing from the quantization side. My Qwen3.8 27B GGUF keeps the MTP head in native NVFP4 on purpose. Upgrading the head to Q5_K/Q6_K (+69 MiB) looked like a free win and instead acceptance went 48.3% to 33.1% and throughput dropped 26.6%. The head only has to agree with the quantized target, not with the BF16 parent, so trimming and quantizing probably win for the same reason. Full numbers: https://piszczek.pl/blog/qwen38-27b-256k-50-tps-24gb-gpu
Michał Piszczek
cdiamond
·
AI & ML interests
CTO of Archdesk · Founder: Lextron.ai, Inclify, Rejsomat.pl, Robotero · Autonomous AI agents · AI reliability · Software architecture · ConstructionTech · Proof-Adjusted Autonomy · Joule Wars
Recent Activity
updated a collection 1 day ago
Qwen3.8 27B at 256K on 24 GB