Comment on Smaller, faster, safer: running Kimi and GLM at scaleparentComments−whimsicalism1modon't disagree, but there is a big difference between 'quantized model/weights' and quantized activations
Comments
don't disagree, but there is a big difference between 'quantized model/weights' and quantized activations