Skip to content

Comment on MLC-LLM: GPT/Llama on consumer-class GPUs and phonesparent

Comments

I find, since 30B models are actually quite usable locally if you have good hardware, that I really want something like a Vicuna 30B. That would be amazing. I can only run 65B locally at a speed of 1 token per second, which is too slow to be usable unfortunately.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.