I think Descript is an incredible product, but as an early beta user of the Lyrebird API, it really frustrates me that this system is currently inaccessible. It was way ahead of its time, and it worked better than anything I've seen since.
I wish that they would release it in some capacity. Heck, even with a licensing fee. There's so many use cases beyond what Descript uses the tech for.
I'm about to release a product on https://FakeYou.com that does real time voice conversion to a target voice. The quality is so-so, but I think we'll improve it quickly.
You can see demos of our voice conversion on https://storyteller.io near the middle of the page (section "3"), where my voice is converted into Donald Trump's voice. (I know, I should have used SpongeBob. We're going to have better product demos soon.)
(We're hiring if you're interested in virtual production. VTubing, Hollywood deepfakes and production pipeline inversion, or even SasS marketing tools.)
What repo is your voice-conversion based on? AutoVC or something like that? The samples page for NANSY and NANSY++ sound really good but I haven't been able to replicate them on my local dev environment just yet.
Would be super interesting if we could create a GAN-based version of this for audio calls that could change your voice into a white male or whatever is necessary for the person on the other end to not impart negative biases on their hiring/investment/business decisions based on your voice.
If they accept you for a job and then discover only later that you're actually female or non-white it'd be a big red flag if they rejected you only after seeing your face.
If someone shows up sounding completely different than they did in the interview then I think most people would assume that the candidate had someone else take the interview for them, resulting in a massive red flag.
People have a rather poor memory and recognition of voices they're not
familiar with. By the time it's been through cellphone codecs (linear
predictive coding that basically resynthesises your voice using a
handful of parameters) it's a wonder we can tell one speaker from
another.
How we do it is to listen for other features of speech, accent and
vowels, speed, rhythm, prosody, intonation, anomalies like glottal
noise, dropped H's, nasal formant, murmuring diphongs.. the things
that make us unique.
Impressionists (mimics) learn those. If they're not present in the
source signal it's not easy to change or add them unless you move the
whole signal to an intermediate form (speech to text) and then
resynthesis the whole show (TTS) via a full articulation model that
has those anomalous features.
If you get any good at this the people who you will piss-off are banks
and folks who use voice as a blind authenticator (hint: your bank
already does if you call them).
Sure, but if the interview was about actual qualifications and skills, then it should not matter.
Even better, companies that were truly dedicated to diversity in hiring could remove names from any materials shown to interviewers and tell interviewees to use the same voice changer.
I wouldn't trust someone who started a relationship with deception. I would have no way of knowing the person is telling the truth that they changed their voice instead of having someone else interview.
The "having someone else interview" problem is vanishingly small compared to racism against minority groups in hiring decisions. That's the real, widespread problem.
It's also pretty easy to figure out within the first week of employment if the person knows their shit or not.
It's not worth going through the effort of onboarding someone for a week to determine if things might work out. Discussions are more than words. Inflection and presentation are all pieces of effective communication. If someone is advocating a voice changer, why even do voice at all? Why not meet in a discord or slack space and type it out?
> Would be super interesting if we could create a GAN-based version of this for audio calls that could change your voice into a white male or whatever is necessary for the person on the other end to not impart negative biases on their hiring/investment/business decisions based on your voice.
Would only work for male -> female (and vice versa) changes. There is no vocal difference between a white male and a black male (or white female and black female); the aural difference is in the accent and actual regional slang, not in the voice.
Being able to modify accents would be really cool though. Differing prosody and regional turns of phrase would likely make the first prototypes sound very uncanny-valley.
Sure, I’m nowhere near a bank director on the list of “attractive targets for voice cloning,” but who knows how widespread this attack might become in the future, and by the time one’s voice is out there on the Net, there’s no way to take it back.
I would like to use a voice changer when making phone calls to businesses too. I can totally imagine future corporations creating a new revenue stream by selling models of a known person’s voice to advertisers so they can later correlate the voice to that person.
Unfortunate that Lyrebird’s transformations seem easy to reverse. I wonder if there are any FOSS tools that make it harder to recover the original voice.
My bank requests voiceprint verification. I personally think it's a terrible idea - I'm against all kinds of 2FA that involve phones or biometrics, and design my own systems to avoid them. I imagine deepfaked voiceprints might be a very lucrative target for thieves in the near future. In the current moment, voice verification is actually an incentive for kidnappers.
Your post gave me an interesting business concept. What if you could buy an anonymous, but totally unique voice, with its own speech patterns and pitches and everything, the way you can currently rent an anonymous email address? Each voice you buy could be generated as a downloadable data set, a one-time sale, subscriptions paid for updates to the software that interprets and reads it.
Unfortunately, a voice changer does not help with that. Human voice and speech is very complex. Just fiddling with some pitches is not enough to fully disguise all that.
Hmmm. I'm not sure that's true. I doubt modern voice-print analyzers are anywhere near sophisticated enough to be able to tease out timing and pronunciation if you've modified the sound beyond a simple pitch shift.
I used it a lot in multiplayer game lobbies to spin elaborate over-the-mic dramas involving a deep-voiced man and his anal-retentive mom. Makes for great fun if you can get your friends in on the action!
I contend that OP's usage of Freudian stage theory in a comic and mischievous scenario is an excellent usage of "debunked pseudoscience" that should be avoided in scenarios where it matters.
If the joke lands because "we all" understand this debunked pseudoscience, what's the harm? If it doesn't land because no one understands Freud anymore, that makes OP a bad comedian.
“I decided to write this as a tool for myself” and presumably the author uses Linux. Seeing as it’s permissively licensed I doubt they’d interfere if someone wanted to port it to other platforms.
https://www.descript.com/lyrebird
https://www.descript.com/overdub