Comment on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/sComments−c0rruptbytes11dso many inference project, omlx already supports all of this and has a 1000 people trying to optimize it constantly−carloslfuOP11dBoth projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX−carloslfuOP11dInteresting! I'll check it out
Comments
so many inference project, omlx already supports all of this and has a 1000 people trying to optimize it constantly
Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX
Interesting! I'll check it out