Skip to content

Comment on Apples to Apples: MLX vs. Llama.cpp for Gemma 4 12B on an M1 16GB

Comments

ABSOP

lol, it took me 48 hours to do (and re-do, and re-do) this test + write it up and now that I convinced myself to stop changing bits and just publish it... Google's just announced the Gemma 4 QAT models :-D

It would not change the core of my article since the bottleneck remains the memory bandwidth on the old M1 16GB though

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.