DeepL Launches Voice-to-Voice Translation Suite
DeepL, famous for its text translation services, unveiled a voice-to-voice translation system on April 16, 2026. This suite handles scenarios such as meetings, conversations on mobile devices and the web, and group discussions tailored for frontline workers using custom applications. The firm also launched an API. This allows external developers and companies to integrate DeepL's technology into their own solutions, including those for call centers.
DeepL started as a German startup in Cologne in 2017. It quickly gained notice for its neural machine translation engine, which many users found more natural and accurate than rivals like Google Translate in several language pairs. The company has expanded to document translation and now steps into voice with this release.
CEO Explains the Move to Voice
"After spending so many years in text translation, voice was a natural step for us," DeepL CEO Jarek Kutylowski said in an interview with TechCrunch. "We have come a long way when it comes to text translation and document translation. But we thought there wasn't a great product for real-time voice translation."
Kutylowski pointed out key hurdles in building real-time voice translation. Developers must balance low latency, the time from speech input to translated output, with high accuracy. DeepL controls its full voice processing pipeline. Right now, it turns speech into text, translates the text, then generates new speech from it. The company credits years of text work for superior translation quality. Plans call for an end-to-end model that handles voice directly, without the text middle step.
Features for Meetings and Conversations
DeepL offers add-ons for Zoom and Microsoft Teams. Users can listen to live translations as speakers use their native languages. They can also view real-time translated text on screen. This early access program has a waitlist for organizations to join.
A separate product supports mobile and web conversations. These work for both in-person and remote interactions. For groups, like in training or workshops, participants scan a QR code to join. The system learns custom vocabulary too. It adapts to industry jargon, company names, and personal names.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Kutylowski sees AI reshaping customer service. A translation layer lets firms offer help in languages short on skilled, affordable workers.
Competition in Voice Translation
DeepL enters a field with established players. Sanas raised $65 million last year from Quadrille Capital and Teleperformance. It uses AI to alter accents live, mainly for call center agents.
Camb.AI, based in Dubai, specializes in speech synthesis and translation for media and entertainment. It partners with Amazon Web Services to dub and localize videos at large scale.
Palabra, supported by Reddit co-founder Alexis Ohanian's Seven Seven Six fund, develops real-time speech translation. It keeps the original meaning and speaker's voice intact, making it a close rival to DeepL's new offering.
Future Plans and Tech Edge
DeepL aims to refine its voice tools over time. The current pipeline relies on proven text strengths. Skipping text-to-speech steps entirely could further drop delays and boost natural flow. With this launch, DeepL positions itself as a full-stack player in real-time communication across languages.

