Skip to content

Comment on Smaller, faster, safer: running Kimi and GLM at scaleparent

Comments

They made an extremely strong claim:

None of this would matter if it changed the model's answers

If they want to assert that the answers don’t change, then perhaps they should calculate the statistical distance between the token probability outputs or something to that effect. I doubt the results would indicate that the answers don’t change by any reasonable interpretation.

Maybe the results are still good enough.

KL divergence is your friend when it comes to evaluating the effects of quantisation: https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_diver...

Isn’t that already in detail by the research of these quantization techniques?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.