Skip to content

Comment on BASE TTS: The largest text-to-speech model to-dateparent

Comments

I've seen a few, there was even one posted to HN some time ago, though I don't recall the exact name. They were working on adding emotion to audio generation, but it was still a bit wonky. Emotion is a tricky concept and one of the reasons (I think) we haven't see a Paul Ekman microexpression detector yet. That's where my suggestion about looking to use action words comes into play, since those are more tangible, offer direction, without trying to identify various emotional valence levels.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.