DeepZen's synthetic AI voices finally make audio books you want to listen to

2 0

By Matthew Griffin Intelligence and the Senses 22nd January 2021

WHY THIS MATTERS IN BRIEF

AI is sounding more and more human-like, and that’s an issue if you make a career out of your voice …

Love the Exponential Future? Join our XPotential Community, future proof yourself with courses from XPotential University, connect, watch a keynote, or browse my blog.

Over the past couple of years synthetic voices, in other words human voices that are generated by Artificial Intelligence (AI) rather than actual people, has come a long way with companies like Facebook and LyreBird being able to replicate anyone’s voice with just a minute’s worth of audio, and tech giants like Google promoting Duplex, which they showed off a while ago, to be the voice of Google Assistant.

Researchers want to revolutionise artificial intelligence by teaching it common sense

As amazing as all these advances are though, and as much as they help push synthetic voice past uncanny valley, the point at which people can’t tell if it’s real or synthetic, and as AI tries to get its head around holding human-like conversations rather than limiting itself to saying a few words or mini-sentences, for the most part none of these systems so far have been able to generate human-like voices that convey emotion, let alone generate a human voice that you’d be happy listening to for hours on end.

DeepZen · She Chose Me by Tracey Emerson (Alexia)

DeepZen · DeepZen Samples

Enter DeepZen, a voice synthesiser project that uses AI algorithms from IBM’s Power AI and Watson technologies division. DeepZen has developed text-to-speech tools that not only sound human at first listen, but can also pick up on the emotional cues needed for reading text in a compelling manner. In doing so, the company claims that it could reduce the time and cost to produce audiobooks by up to 90 percent.

Baxter the robot fixes its mistakes by reading your mind

Taylan Kamis, CEO and Co-founder of DeepZen, explains: “Our aim isn’t to put voice actors out of jobs, but rather to solve the capacity issues in the current market. We identify emotion in text automatically and use voice samples – for which we pay royalties to voice actors – combined with speech synthesis technology to produce convincing voice audio.”

“To [create these voices], we needed to create large and complex neural networks. These require extensive amounts of processing power to produce accurate results fast, so we needed the right technology platform to bring our vision to life,” he added.

US legislators new SELF DRIVE Act lays the foundation for a country full of driverless vehicles

While DeepZen promises it’s not going to put narrators out of a job though it’s hard to see how that won’t happen over the longer term, but that conversation asides in the interim their technology will no doubt be an invaluable tool in helping smaller publishers and indie authors create audiobooks without having to go to the hassle of dealing with professional narrators.

And as for next steps, DeepZen have announced that they’re going to be working with Audiowhale to commercialise their technology and bring it to authors and publishers everywhere.

Matthew Griffin / About Author

Matthew Griffin, multi-award winning Futurist and named Futurist of the Year 2024, has been described as a "Walking encyclopaedia of the future" by NASA and a futurist polymath. One of the world's most renowned futurists and strategic foresight experts Matthew is the 15 times author of the blockbuster "Codex of the Future" series, and is the Founder and Futurist in Chief of the 311 Institute, a global Futures and Deep Futures advisory firm working across the next 50 years, XPotential University, the world's first free futures and foresight university, and the World Futures Forum which works with the United Nations to solve the worlds greatest challenges. Matthew is an in demand international keynote, acclaimed university lecturer and mentor, and host of the hit Fanatical Futurist podcast.

A rare talent in his past Matthew helped build and run several multi-billion dollar business units for Atos, Dell-EMC, and IBM, and his ability to identify, track, and explain the impacts of hundreds of emerging technologies and trends on global business, culture, and society has earned him a powerful reputation and a roster of clients that include royal households, world leaders, G7, G20, and G77+ governments, and many of the world's most respected brands including ABB, Accenture, Adidas, AON, ARM, BCG, Centrica, Citi Group, Coca Cola, Dentons, Deloitte, Disney, Dow, EY, KPMG, Lego, Legal & General, LinkedIn, Microsoft, PepsiCo, Qualcomm, RWE, Samsung, T-Mobile, UBS, VISA, and many others. He was also the only futurist invited to talk at the UN COP28 held in Dubai alongside world leaders.

Regularly featured in the global media including the AP, BBC, Bloomberg, CNBC, Discovery, Forbes, Khaleej Times, Telegraph, TIME, ViacomCBS, WIRED, and the WSJ, Matthews mission is to help organisations create a fair and sustainable future whose benefits are shared by everyone irrespective of their ability, background, or circumstances.