Comment on MLC-LLM: GPT/Llama on consumer-class GPUs and phonesparentComments−killthebuddha3yThis is not a direct answer to your question, but performance is better in terms of the _quality_ of completions but not in terms of price, latency, or uptime.
Comments
This is not a direct answer to your question, but performance is better in terms of the _quality_ of completions but not in terms of price, latency, or uptime.