Skip to content

Comment on Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

Comments

Wow, software is hard! Imagine an entire company working to build an insanely huge and expensive wafer scale chip and your super smart and highly motivated machine learning engineers get 1/3 of peak performance on their first attempt. When people say NVIDIA has no moat I'm going to remember this - partly because it does show that they do, and partly because it shows that with time the moat can probably be crossed...

make it work, make it work right(ish), now make it fast.

Fast and wrong is easy!

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.