In January 2020, I have launched a website with a conversational AI for language learning.
Conversational artificial intelligence (AI) refers to technologies, like chatbots, which users can exchange messages with.
With the first version of the website, the users can only send written messages.
I have received feedback from users who want to listen to the messages, so I have started to investigate on the available APIs for text-to-speech.
I have realized that the AI voices were getting really close to natural human voices.
The biggest problem for me is that the AI will fail to read correctly some unusual words, so you still have to double check the result to be 100% sure of the result (If you hire a voice-over talent on a freelance platform, you still have to double check the result anyway).
You can find comments and articles online about people hating artificial voices, like in this article about TikTok synthetic voices.
https://beyondwords.io/blog/heres-why-people-hate-tiktoks-text-to-speech-voices/
"TikTok's synthetic voices sound robotic, too. Even if you ignore mispronunciations, their tone and inflections often sound very unnatural."
I think that some people hate artificial voices because the first time they have listened to a synthetic voice, it was an inferior one.
We are already at the point where the best AI voices are sounding almost natural, especially for the non native speakers.
With AI voices, most people can now produce audios and videos voice-overs, without having a proper place and audio equipment to record their content.
Right now for me it would be impossible to record an audio because a neighbor is blasting music on his radio, and my kids can randomly start to shout at each other because of their petty fight.
Of course the AI voices will never compete with the best voice-over talents. Human voice-overs and AI voice-overs should coexist without problem.