Skip to content

Comment on H3-metal – Native MiniMax-H3 inference for Apple Silicon

Comments

This is where the DGX spark makes up a bit of the ground it loses on llm work, diffusion and cuda go together like peanut butter and jelly.

cough DiffusionGemma cough

Seriously, very dumb model compared to what you can run locally, but holy moly is it FAST on one GPU, seriously impressive. Can't wait for those to be scaled up a bit to fit perfectly within 96GB VRAM, then they'll be competitive.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.