Skip to content

Comment on Mesh LLM: distributed AI computing on irohparent

Comments

I’m staring at this comment for a while now: With 3ms latency combined per token, wouldn’t that mean (1 / latency) = 333 token/s for the theoretical upper bound? I’m not trying to nitpick, just curious if I misunderstand something.

Indeed, I completely screwed my math up. Looks like 10am is too early in the morning for a Sunday.

33 tps max token generation speed would be for 10ms of network latency in the above example.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.