Comment on Smaller, faster, safer: running Kimi and GLM at scaleparentComments−bigbaguette1moCould you suggest good resources about LLM serving? Blogs, articles...−vietvu1moI cannot think of any that so concentrated. Best place to me is r/localllama on reddit. Unsloth website and twitter posts is a good place too (usually in localllama too).−Tepix1moThe latest Latent.Space episode https://www.latent.space/p/inference-eng
Comments
Could you suggest good resources about LLM serving? Blogs, articles...
I cannot think of any that so concentrated. Best place to me is r/localllama on reddit. Unsloth website and twitter posts is a good place too (usually in localllama too).
The latest Latent.Space episode https://www.latent.space/p/inference-eng