Skip to content

Comment on H3-metal – Native MiniMax-H3 inference for Apple Silicon

Comments

On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.

Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway

On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.

That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.

It’s always been the case, it’s more the anomaly that LLMs work at comparable speeds on M series because almost all other ML runs way faster on Nvidia cards.

LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.

An RTX6000 is a completely different class of hardware.

Really? No wonder I keep trying to type on it like a laptop but it doesn't work and doesn't even have a display!

Right? Do you also type on a Mac Studio without a keyboard plugged in? Like tap the ethernet port 3 times in a row then this sequence of sticking your fingers into TB5 ports? I mean, it’s clear you stick your RTX into a computer.

Please keep up posted about the results!

First batch of quick test results: approximately 1/5 speed improvement

+20% or x5 speed improvement?

Put Codex to work on deploying it now

Which codex?

For gods sakes, Apple let people have run other GPUs instead of these pissweak 2012 class mobile GPUs

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.