Skip to content

Comment on Fable 5 pushed Gemma 4 to 255 tok/s on WebGPUparent

Comments

Either ollama or omlx, both are pretty dang performant. Omlx lets you run Claude code locally though as long as you bootstrap it with the right model

Omlx is really nice, thanks for the recommendation!

Why would you need Omlx? For speed up?

Has extra KV cache on SSD, and lots more options to tweak. There's experimental TurboQuant and multi token prediction support.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.