This page was machine-translated and may differ from the original. View original

DeepL Launches Real-time Voice Translation with Speaker Characteristics Preservation

Google 우선 소스Published2026.09.16 12:19


 
New Voice AI Model and Desktop App Launch, Supporting Voice-to-Voice Translation in Over 30 Languages
 
DeepL has expanded the scope of its voice translation technology from word delivery to preservation of speaker-specific speech characteristics. With this update, voice elements that distinguish speakers—such as speaking style, rhythm, speed, and intonation—are maintained in translation results, enabling more natural communication in multilingual real-time conversation environments.
 
DeepL unveiled a major update to its voice AI technology "DeepL Voice" on the 16th.

The new voice AI model reproduces each speaker's unique speech characteristics in translated voice, supporting voice preservation in 14 languages including Korean and real-time voice-to-voice translation in over 30 languages.
 
The core of this update is speaker identification-based voice preservation technology.

Even in environments where multiple participants speak simultaneously, the technology distinguishes and maintains individual speaker characteristics.

Key features include △voice preservation △natural expression delivery △low-latency real-time translation △support for voice-to-voice translation in over 30 languages.

The natural expression delivery feature reflects expression elements conveyed in speech—such as question nuances, hesitations, and urgency—in the translated speech.
 
Sebastian Enderlein, Chief Technology Officer (CTO) of DeepL, stated, "We faced a technical challenge of maintaining the unique identity of voice while ensuring both translation quality and low latency," and added, "This will open new possibilities extending from corporate operations and customer service to multilingual product development."
 
DeepL has also launched a new voice desktop app available on Windows and Mac environments.

The app integrates with major online meeting platforms including Zoom, Microsoft Teams, and Google Meet, providing a consistent translation environment without additional configuration.

It automatically detects calls and displays real-time translation subtitles in a constant overlay format, and supports switching between subtitle and voice translation, as well as adjustments to subtitle size, audio volume, and translation settings.

The new voice features can be used directly in the app without depending on platform-specific updates.

Jarek Kutylowski, founder and CEO of DeepL, stated, "DeepL Voice is evolving into an audio-based AI system that preserves not only subtitle translation but also user identity, emotion, and speaking timing in real-time."
To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.
명세환 기자
명세환 Reporter

Comments