Pool spare GPU capacity to run LLMs at larger scalegithub.com/michaelneale 11 pointsi3865 months ago3 commentsSaveHideCopy link On HNComments−lostmsu5moMoE models via expert sharding with zero cross-node inference trafficThis makes the whole project questionable−vagrantJin5moThis is very promising, definitely looks more user friendly than exo. Can't wait to try it out.−iwinux5moYou lost me on "spare GPU". I don't have any capable GPUs, let alone spare ones :)
Comments
This makes the whole project questionable
This is very promising, definitely looks more user friendly than exo. Can't wait to try it out.
You lost me on "spare GPU". I don't have any capable GPUs, let alone spare ones :)