Skip to content

Comment on Ollama: All Aboard Open Models

Comments

https://github.com/ollama/ollama/issues/11772

A year and still no implementation for such a basic need as offloading MoE layers onto the CPU selectively. On llama.cpp I can get models like Qwen 35BA3B running partially on gpu/cpu with 40t/s on a laptop thanks to --n-cpu-moe but on this VC funded joke it would be simply unusable. I can't quite understand how you make a wrapper so much worse than the code you're ripping out.

This funding is fuel for what’s ahead. Ollama sits front and center in the open model ecosystem

No.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.