The news: With many of us now relying on video calls for face-to-face interaction, choppy connections are more frustrating than ever. An artificial intelligence that mimics an individual speaker’s way of talking can smooth over the cracks by filling in small gaps with snippets of generated speech. Developed by a team at Google, the technology is now being used in Google’s video-calling app Duo.
What’s the problem? When you’re on an online call your voice gets chopped up into lots of tiny pieces that are zipped across the internet in data blocks known as packets. Packets often arrive at the other end jumbled up and software has to reorder them. But sometimes packets don’t arrive at all, which creates glitches and gaps in a conversation. This happens at the best of times. According to Google 99% of Duo calls have to deal with jumbled up or lost packets. A tenth of those calls lose more than 8% of their audio.
Generating speech: To fix the problem, the team built on a neural network developed by DeepMind that can generate realistic speech from text. Called WaveNetEQ, the new neural network was then trained on a large dataset of 100 recorded human voices speaking 48 different languages until it could auto-complete short sections of speech based on common patterns in the way people talk. Because Duo is end-to-end encrypted, the AI runs on the device, not the cloud. During a call, WaveNetEQ is able to learn characteristics of a speaker’s voice and generates audio snippets that match both the style and content of what the speaker is saying. When a packet is lost, the AI generated voice is inserted in its place.
For now, the AI can only generate syllables rather than whole words or phrases. But short samples Google posted online show that the results can be pretty lifelike. In one case, the AI replaces the second syllable of the word “trouble” in a voice that mimics the male speaker exactly.
This artist is dominating AI-generated art. And he’s not happy about it.
Greg Rutkowski is a more popular prompt than Picasso.
What does GPT-3 “know” about me?
Large language models are trained on troves of personal data hoovered from the internet. So I wanted to know: What does it have on me?
DeepMind has predicted the structure of almost every protein known to science
And it’s giving the data away for free, which could spur new scientific discoveries.
An AI that can design new proteins could help unlock new cures and materials
The machine-learning tool could help researchers discover entirely new proteins not yet known to science.
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.