Skip to content

Comment on Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

Comments

I wonder if there is a token/watt metric. Afaiu cerebras uses plenty of power/cooling.

I found this on their product page, though just for peak power:

At 16 RU, and peak sustained system power of 23kW, the CS-3 packs the performance of a room full of servers into a single unit the size of a dorm room mini-fridge.

It's pretty impressive looking hardware.

https://cerebras.ai/product-system/

Weighing 800kg (!). Like, what the heck.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.