A collection of reproducible LLM inference engine benchmarks: SGLang vs. vLLMgithub.com/Michaelvll 1zhwu1ydiscuss
Efficient GPU Resource Management for ML Workloads Using SkyPilot, Kueue on GKEgithub.com/GoogleCloudPlatform 2zhwu1ydiscuss
New Recipe: Serving Llama-2 with VLLM's OpenAI-Compatible API Servergithub.com/skypilot-org 1zhwu3ydiscuss
Vicuna releases its secrete of finding available A100s on the cloud to train ittwitter.com 4zhwu3y2 comments