Skip to content

Comment on Smaller, faster, safer: running Kimi and GLM at scale

Comments

Thanks for the transpiration but this is too shallow when talking about LLM serving.

Could you suggest good resources about LLM serving? Blogs, articles...

I cannot think of any that so concentrated. Best place to me is r/localllama on reddit. Unsloth website and twitter posts is a good place too (usually in localllama too).

The latest Latent.Space episode https://www.latent.space/p/inference-eng

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.