Skip to content

Comment on Smaller, faster, safer: running Kimi and GLM at scaleparent

Comments

Worth knowing a cache hit rate can be structurally zero and look identical to a broken one.

Of course. On the other hand if you see 8 % gap for similar workloads, averaged across tens of sessions, with the same underlying model, it becomes a pretty clear signal. And I do exclude first request per provider per session from the statistic.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.