Comment on Deep Voice: Real-Time Neural Text-To-SpeechparentComments−Dowwie9ynow? https://www.youtube.com/watch?v=XfcqBElF0ZI−M4v3R9yAfaik VoCo isn't creating anything from thin air, instead it scans the available voice data (it reportedly needs a sample of about 20 mins of a person speaking) and copies fragments of it in specific order to create a sentence.
Comments
now? https://www.youtube.com/watch?v=XfcqBElF0ZI
Afaik VoCo isn't creating anything from thin air, instead it scans the available voice data (it reportedly needs a sample of about 20 mins of a person speaking) and copies fragments of it in specific order to create a sentence.