I actually didn't try to mask further, but you made it interesting to try.
The published acceptance scalar is accepted/drafted. It also logs accepted tokens per verify round and conditional acceptance by draft position. So your second decomposition is the matching one for my numbers: roughly 6.6–7.7% step-cost saving, rather than the 8.25% i.i.d. estimate. I also agree the DFlash2 M5 Max result is not directly portable to my RTX Blackwell/Memra/NVFP4 setup.
I did try DFlash2, but with my rig it doesn't reproduce the results mentioned; the gains are
Workload MTP K=3 DFlash2 DFlash2 delta
16 chat prompts, steady 78.82 tok/s 80.51 tok/s +2.1%
16 agentic prompts, steady 77.81 tok/s 81.66 tok/s +4.9%
Nice improvement, but there's something else that also affect the numbers.