Comment on H3-metal – Native MiniMax-H3 inference for Apple SiliconparentComments−whywhywhywhy1moIt’s always been the case, it’s more the anomaly that LLMs work at comparable speeds on M series because almost all other ML runs way faster on Nvidia cards.
Comments
It’s always been the case, it’s more the anomaly that LLMs work at comparable speeds on M series because almost all other ML runs way faster on Nvidia cards.