Comment on Cerebras launches inference for Llama 3.1; benchmarked at 1846 tokens/s on 8BparentComments−twothreeone2yNot if you're serving tens of thousands of users at the same time.−cma2yStill tiny at 100,000.
Comments
Not if you're serving tens of thousands of users at the same time.
Still tiny at 100,000.