Comment on Mesh LLM: distributed AI computing on irohparentComments−ul52552moI’m staring at this comment for a while now: With 3ms latency combined per token, wouldn’t that mean (1 / latency) = 333 token/s for the theoretical upper bound? I’m not trying to nitpick, just curious if I misunderstand something.−stymaar2moIndeed, I completely screwed my math up. Looks like 10am is too early in the morning for a Sunday.33 tps max token generation speed would be for 10ms of network latency in the above example.
Comments
I’m staring at this comment for a while now: With 3ms latency combined per token, wouldn’t that mean (1 / latency) = 333 token/s for the theoretical upper bound? I’m not trying to nitpick, just curious if I misunderstand something.
Indeed, I completely screwed my math up. Looks like 10am is too early in the morning for a Sunday.
33 tps max token generation speed would be for 10ms of network latency in the above example.