Comment on Fable 5 pushed Gemma 4 to 255 tok/s on WebGPUparentComments−thomspoon2moEither ollama or omlx, both are pretty dang performant. Omlx lets you run Claude code locally though as long as you bootstrap it with the right model−mike_hearn2moOmlx is really nice, thanks for the recommendation!−karussell2moWhy would you need Omlx? For speed up?−pornel2moHas extra KV cache on SSD, and lots more options to tweak. There's experimental TurboQuant and multi token prediction support.
Comments
Either ollama or omlx, both are pretty dang performant. Omlx lets you run Claude code locally though as long as you bootstrap it with the right model
Omlx is really nice, thanks for the recommendation!
Why would you need Omlx? For speed up?
Has extra KV cache on SSD, and lots more options to tweak. There's experimental TurboQuant and multi token prediction support.