Comment on Smaller, faster, safer: running Kimi and GLM at scaleComments−vietvu1moThanks for the transpiration but this is too shallow when talking about LLM serving.−bigbaguette1moCould you suggest good resources about LLM serving? Blogs, articles...−vietvu1moI cannot think of any that so concentrated. Best place to me is r/localllama on reddit. Unsloth website and twitter posts is a good place too (usually in localllama too).−Tepix1moThe latest Latent.Space episode https://www.latent.space/p/inference-eng
Comments
Thanks for the transpiration but this is too shallow when talking about LLM serving.
Could you suggest good resources about LLM serving? Blogs, articles...
I cannot think of any that so concentrated. Best place to me is r/localllama on reddit. Unsloth website and twitter posts is a good place too (usually in localllama too).
The latest Latent.Space episode https://www.latent.space/p/inference-eng