Skip to content

Comment on BASE TTS: The largest text-to-speech model to-dateparent

Comments

XTTS has a streaming mode with ~300ms latency and sounds good, though it has hallucination issues. StyleTTS2 sounds good and doesn't hallucinate as much. It doesn't support streaming but it's fast so it can still respond quickly. But neither of them sound as good as Eleven Labs or OpenAI or this one.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.