We use vllm as it generally has the best ecosystem support.
Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit patterns).
We've never had a limitation at the tokenizer step. Limitations at peak tend to manifest more on slower time in vllm doing prefill or decode though we actively try and minimize this.
Comments
Whats the prefered LLM runtime to use? vLLM?
Tips & Tricks on parameters/settings?
What happens at peak? Do people have to wait now? Increase of latency?
We use vllm as it generally has the best ecosystem support. Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit patterns). We've never had a limitation at the tokenizer step. Limitations at peak tend to manifest more on slower time in vllm doing prefill or decode though we actively try and minimize this.