Skip to content

Comment on Smaller, faster, safer: running Kimi and GLM at scaleparent

Comments

That’s disappointing, in the past cloudflare had some of the best engineering blog articles

I don't disagree, but at some point in the last year they ended up severely word-expanded. So in a revealed sense, they are no longer meant for human consumption except for those who don't significantly value their own time. There is very little information in the post that an agent can't pull for you:

* they use quantized models

* they quantize KV cache

* they have a cache tagging mechanism to prevent cache misuse (neat)

The agent can extract numbers without filler prose as well.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.