Comment on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/sparentComments−nixon_why6910dI find mtp=3 does well with that model, only at 4 it becomes unprofitable.Check your quants, its worth having the mtp layer be a bigger quant if it leads to 2x throughput from more accepted tokens.
Comments
I find mtp=3 does well with that model, only at 4 it becomes unprofitable.
Check your quants, its worth having the mtp layer be a bigger quant if it leads to 2x throughput from more accepted tokens.