TensorRT-LLM runtime now open-sourcegithub.com/NVIDIA 4 pointsmmoskal1 year ago1 commentSaveHideCopy link On HNComments−mmoskalOP1yPreviously, the "Executor" runtime was shipped as binary blobs. This is the bit that schedules requests and manages KV cache (similar to vLLM or SGLang server).
Comments
Previously, the "Executor" runtime was shipped as binary blobs. This is the bit that schedules requests and manages KV cache (similar to vLLM or SGLang server).