The news: With many of us now relying on video calls for face-to-face interaction, choppy connections are more frustrating than ever. An artificial intelligence that mimics an individual speaker’s way of talking can smooth over the cracks by filling in small gaps with snippets of generated speech. Developed by a team at Google, the technology is now being used in Google’s video-calling app Duo.
What’s the problem? When you’re on an online call your voice gets chopped up into lots of tiny pieces that are zipped across the internet in data blocks known as packets. Packets often arrive at the other end jumbled up and software has to reorder them. But sometimes packets don’t arrive at all, which creates glitches and gaps in a conversation. This happens at the best of times. According to Google 99% of Duo calls have to deal with jumbled up or lost packets. A tenth of those calls lose more than 8% of their audio.
Generating speech: To fix the problem, the team built on a neural network developed by DeepMind that can generate realistic speech from text. Called WaveNetEQ, the new neural network was then trained on a large dataset of 100 recorded human voices speaking 48 different languages until it could auto-complete short sections of speech based on common patterns in the way people talk. Because Duo is end-to-end encrypted, the AI runs on the device, not the cloud. During a call, WaveNetEQ is able to learn characteristics of a speaker’s voice and generates audio snippets that match both the style and content of what the speaker is saying. When a packet is lost, the AI generated voice is inserted in its place.
For now, the AI can only generate syllables rather than whole words or phrases. But short samples Google posted online show that the results can be pretty lifelike. In one case, the AI replaces the second syllable of the word “trouble” in a voice that mimics the male speaker exactly.
This new data poisoning tool lets artists fight back against generative AI
The tool, called Nightshade, messes up training data in ways that could cause serious damage to image-generating AI models.
Rogue superintelligence and merging with machines: Inside the mind of OpenAI’s chief scientist
An exclusive conversation with Ilya Sutskever on his fears for the future of AI and why they’ve made him change the focus of his life’s work.
Driving companywide efficiencies with AI
Advanced AI and ML capabilities revolutionize how administrative and operations tasks are done.
Unpacking the hype around OpenAI’s rumored new Q* model
If OpenAI's new model can solve grade-school math, it could pave the way for more powerful systems.
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.