Comment on Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/sComments−neals1ySo what is inference?−jonplackett1yInference just means using the model, rather than training it.As far as I know Nvidia still has a monopoly on the training part.
Comments
So what is inference?
Inference just means using the model, rather than training it.
As far as I know Nvidia still has a monopoly on the training part.