Spotify is going to clone podcasters’ voices — and translate them to other languages

stopthatgirl7@kbin.social · 1 year ago

Spotify is going to clone podcasters’ voices — and translate them to other languages

FireWire400@lemmy.world · 1 year ago

That’s just weird… Part of the reason I listen to podcasts is that I just enjoy people talking about things and AI voices still have this uncanny quality to me

danielbln@lemmy.world · 1 year ago

http://sndup.net/tb9p

nehal3m@sh.itjust.works · 1 year ago

Point taken, well done.

Hoimo@ani.social · 1 year ago

That’s obviously way better than any TTS before it, but I still wouldn’t want to listen to it for more than a few minutes. In these two sentences I can already hear some of the “AI quirks” and the longer you listen, the more you start to notice them.
I listen to a lot of AI celeb impersonations and they all sound like the same machine with different voice synthesizers. There’s something about the prosody that gives it away, every sentence has the same generic pattern.
Humans are generally more creative, or more monotonous, but AI is in a weird inbetween space where it’s never interested and never bored, always soulless.

bamboo@lemm.ee · 1 year ago

Having listened to it, I could not identify any sort of “AI quirk”. It sounded perfectly fine.

TwinTusks@outpost.zeuslink.net · 1 year ago

This is beautiful

rigatti@lemmy.world · 1 year ago

It won’t take long until that uncanny quality is worked out.

danielbln@lemmy.world · 1 year ago

Imho it has already been worked out. There is probably selection bias at play as you don’t even recognize the AI voices that are already there.

Pantoffel@feddit.de · 1 year ago

Following up on the other comment.

The issue is that widely available speech models are not yet offering the quality that is technically possible. That is probably why you think we’re not there yet. But we are.

Oh, I’m looking forward to just translate a whole audiobook into my native language and any speaking style I like.

Okay, perhaps we would still have difficulties with made up fantasy words or words from foreign languages with little training data.

Mind, this is already possible. It’s just that I don’t have access to this technology. I sincerely hope that there will be no gatekeeping to the training data, such that we can train such models ourselves.

Spotify is going to clone podcasters’ voices — and translate them to other languages

Spotify is going to clone podcasters’ voices — and translate them to other languages

Spotify is going to clone podcasters’ voices —&nbsp;and translate them to other languages

Spotify is going to clone podcasters’ voices — and translate them to other languages