Skip to content

Comment on Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/sparent

Comments

Whisper runs so well locally on any hardware I’ve thrown at it, why run it in the cloud?

Does it run well on CPU? I've used it locally but only with my high end (consumer/gaming) GPU, and haven't got round to finding out how it does on weaker machines.

It’s not fast but if your transcript doesn’t have to get out ASAP it’s fine

That's pretty much exactly how I started. Ran whisper.cpp locally for a while on a 3070Ti. It worked quite well when n=1.

For our use case, we may get 1 audio file at a time, we may get 10. Of course queuing them is possible but we decided to prioritize speed & reliability over self hosting.

Got it. Makes sense in that context

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.