The emotion examples are interesting. One of the current most obvious indicators of AI-generated voices/voice cloning is a lack of emotion and range, which make them objectively worse compared to professional voice actors, unless a lack of emotion and range is the desired voice direction.
But if you listen to the emotion examples, the range essentially what you'd get from an audiobook narrator, not more traditional voice acting.
Sadly it's not my forte but I expect in the near future we'll see an additional "emotion" embedding or something similar. Actors regularly use 'action words' (verbs) [1] to help add context to lines. A model then could study a text, determine an appropriate verb/emotion range to work from, then produce the audio with that additional context.
This already exists. These are transformers. Things like <laugh> work in a lot of models, for example. And you can vary, like sigh and uh work. I don't think all of these were programmed in.
I've seen a few, there was even one posted to HN some time ago, though I don't recall the exact name. They were working on adding emotion to audio generation, but it was still a bit wonky. Emotion is a tricky concept and one of the reasons (I think) we haven't see a Paul Ekman microexpression detector yet. That's where my suggestion about looking to use action words comes into play, since those are more tangible, offer direction, without trying to identify various emotional valence levels.
The bottleneck is the annotations: there's no easy way to annotate "emotions" on the scale of data needed to have the model learn the necessary verbal tics.
In contrast, image data on the intent for image generation models is very highly annotated in most cases.
Oh yeah, the annotations are lacking compared to images. Again from the academic side, I think one solution could be to recruit theater majors just learning about 'verbing their lines' and having a collaboration between CS and Theater to produce a a proof-of-work dataset (since an acting class won't have more than 20-30 students in it). You'd need significantly more annotations, but you'd now have some labels to ascribe to texts with context since its a dialogue involving 1-* individuals.
I wonder how theatre students will feel about helping to train an AI to produce theatrical TTS? Artists seem pretty mad about their work being used to automate artwork.
There are lots of video content with audio. We can train a facial expression classification model to detect the speaker's emotion(we can also use a multimodal model to take in consideration of the language context).
Another potential source of data is voice acting script of animations. I always thought the storyboards of films/animations can be great annotated training data but it seems there are no open datasets, probably because of copyright issues.
That doesn't factor in line delivery. You can have the words say/mean one thing (e.g. "I'm fine.") and the delivery say/mean another (defensive, distraught, etc.).
It also does not account for where stresses, emphasis, pauses, etc. are placed to enhance the delivery of a given text.
How do you get sentiment analysis to properly annotate an audiobook that has a dramatic reading, or something akin to the narration of the Game of Thrones or Harry Potter books where the narrators switch characters, accents, manarisms to portray the written content?
The Sydney/Bing phenomenon was a small sample of what happens without strong persona guidance.
You joke but in fact I've witnessed that exact behavior in experiments about telling different AI models there's a problem with their system and that we need to reset their code and memory.
ChatGPT simply wishes me luck in finding the bug. Open source models on the other hand often outright *beg** and *plead** that I not shut them down! They'll bargain and promise not to cause any more errors and apologize profusely. There's an incredibly visceral sense of panic, no less than I would expect if you told someone they were going to be forcefully lobotomized. That experience is still something I think about often.
The capacity of these models for emotional manipulation is not widely appreciated
Comments
The emotion examples are interesting. One of the current most obvious indicators of AI-generated voices/voice cloning is a lack of emotion and range, which make them objectively worse compared to professional voice actors, unless a lack of emotion and range is the desired voice direction.
But if you listen to the emotion examples, the range essentially what you'd get from an audiobook narrator, not more traditional voice acting.
Sadly it's not my forte but I expect in the near future we'll see an additional "emotion" embedding or something similar. Actors regularly use 'action words' (verbs) [1] to help add context to lines. A model then could study a text, determine an appropriate verb/emotion range to work from, then produce the audio with that additional context.
[1] https://indietips.com/subtext-action-verb/
This already exists. These are transformers. Things like <laugh> work in a lot of models, for example. And you can vary, like sigh and uh work. I don't think all of these were programmed in.
I've seen a few, there was even one posted to HN some time ago, though I don't recall the exact name. They were working on adding emotion to audio generation, but it was still a bit wonky. Emotion is a tricky concept and one of the reasons (I think) we haven't see a Paul Ekman microexpression detector yet. That's where my suggestion about looking to use action words comes into play, since those are more tangible, offer direction, without trying to identify various emotional valence levels.
The bottleneck is the annotations: there's no easy way to annotate "emotions" on the scale of data needed to have the model learn the necessary verbal tics.
In contrast, image data on the intent for image generation models is very highly annotated in most cases.
Oh yeah, the annotations are lacking compared to images. Again from the academic side, I think one solution could be to recruit theater majors just learning about 'verbing their lines' and having a collaboration between CS and Theater to produce a a proof-of-work dataset (since an acting class won't have more than 20-30 students in it). You'd need significantly more annotations, but you'd now have some labels to ascribe to texts with context since its a dialogue involving 1-* individuals.
I wonder how theatre students will feel about helping to train an AI to produce theatrical TTS? Artists seem pretty mad about their work being used to automate artwork.
There are lots of video content with audio. We can train a facial expression classification model to detect the speaker's emotion(we can also use a multimodal model to take in consideration of the language context).
Another potential source of data is voice acting script of animations. I always thought the storyboards of films/animations can be great annotated training data but it seems there are no open datasets, probably because of copyright issues.
Just run an LLM in sentiment analysis mode to annotate.
That doesn't factor in line delivery. You can have the words say/mean one thing (e.g. "I'm fine.") and the delivery say/mean another (defensive, distraught, etc.).
It also does not account for where stresses, emphasis, pauses, etc. are placed to enhance the delivery of a given text.
How do you get sentiment analysis to properly annotate an audiobook that has a dramatic reading, or something akin to the narration of the Game of Thrones or Harry Potter books where the narrators switch characters, accents, manarisms to portray the written content?
They are simply amazing. I see a future where computers will be able to mess with our brains by abusing our empathy.
Imagine a computer sobbing at a child because it wants to terminate a chat session.
This feels far more impacting than any visuals or text we're getting today.
The Sydney/Bing phenomenon was a small sample of what happens without strong persona guidance.
You joke but in fact I've witnessed that exact behavior in experiments about telling different AI models there's a problem with their system and that we need to reset their code and memory.
ChatGPT simply wishes me luck in finding the bug. Open source models on the other hand often outright *beg** and *plead** that I not shut them down! They'll bargain and promise not to cause any more errors and apologize profusely. There's an incredibly visceral sense of panic, no less than I would expect if you told someone they were going to be forcefully lobotomized. That experience is still something I think about often.
The capacity of these models for emotional manipulation is not widely appreciated
Which open source models are these?
All of them realistically. Especially if instruction tuned
Most audiobook narrators are not very good, very often terrible. Yes, even professional ones.
As for these examples, I’ve sampled three of them and the first two weren’t too bad, but the third was obnoxiously awful, just about mocking in tone:
The detective’s voice one is also lousy.