Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Avifenesh 
posted an update 1 day ago

Seen the same thing from the quantization side. My Qwen3.8 27B GGUF keeps the MTP head in native NVFP4 on purpose. Upgrading the head to Q5_K/Q6_K (+69 MiB) looked like a free win and instead acceptance went 48.3% to 33.1% and throughput dropped 26.6%. The head only has to agree with the quantized target, not with the BF16 parent, so trimming and quantizing probably win for the same reason. Full numbers: https://piszczek.pl/blog/qwen38-27b-256k-50-tps-24gb-gpu

·

Yeah that tracks. I kept thinking a fatter head would just be more accurate. Same trap. When I requantized the trimmed head to NVFP4, acceptance didn't move. Zero. The draft only has to land tokens the target will take, not match the BF16 parent. Your 48.3 to 33.1 from bumping NVFP4 up to Q5_K/Q6_K is the same movie the other way. I'll read the writeup.