Skip to content

Comment on Smaller, faster, safer: running Kimi and GLM at scale

Comments

So they quantize models, only tell about it in the blog post (instead of a warning on the model page), and even in the blog post pretend there's no difference by benchmarking on small context tasks many of which are saturated. Coding agents will probably be severely negatively affected by KV quantization.

I'd say serving quantized models without saying so on the "store" page is fraud.

don't disagree, but there is a big difference between 'quantized model/weights' and quantized activations

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.