Skip to content

Comment on Hertz-dev, the first open-source base model for conversational audio

Comments

Cool, looks like this is trained on 16 million hours of audio (500B tokens at ~.11 seconds per token).

Even the large open source TTS models (see F5 TTS, Mask GCT) are mostly trained on very small audio datasets (say 100k hours) relative to the amount of audio available on the internet, so it's cool to see an open source effort to scale up training significantly.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.