Alibaba Qwen3.8 LiveTranslate Drops Speech Lag to 2.3s
Alibaba has launched Qwen3.8 LiveTranslate, a simultaneous interpretation model that reduces translation lag to 2.3 seconds and preserves individual voices across 60 input languages.

Alibaba has introduced Qwen3.8 LiveTranslate, a hosted, closed-weight model designed for real-time speech translation. Built on a novel Interleave architecture, the model processes incoming media and outputs translated text or speech within a single system, bypassing the traditional chain of separate speech recognition, translation, and synthesis models. This update reduces length-adaptive average lagging from 2.8 seconds in the previous Qwen3.5 version to 2.3 seconds. It supports 60 input languages and 29 spoken output languages, maintaining over 94 percent of offline translation quality.
The model introduces real-time speaker diarization to distinguish between multiple voices in a single stream. It then uses voice cloning to preserve the unique vocal characteristics of each speaker in the translated audio. To ensure consistency across long sessions, Qwen3.8 LiveTranslate utilizes its 53,000-token context window to track conversation history, preventing the mistranslation of technical jargon, product terms, and names. It can also ingest video frames at 0.5 tokens per 28-by-28-pixel patch, allowing visual cues like lip movements and gestures to inform the translation.
Available via a WebSocket API on Alibaba Cloud Model Studio, the model is priced at $7.50 per one million audio input tokens. Audio consumption is billed at 12.5 tokens per second for both input and output, which translates to roughly $0.0281 per minute of continuous bilingual audio, or about $1.69 per hour. Default rate limits are set at 100,000 tokens per minute and 10 requests per minute. While earlier evaluations of the Flash line placed Qwen ahead of Gemini 2.5 Flash, GPT-4o Audio Preview, and Voxtral Small 24B in translation accuracy, those tests do not reflect the new version's latency and diarization improvements.
For developers, this single-model approach eliminates the complexity of orchestrating separate speech-to-text, translation, and text-to-speech APIs. However, because the model is closed-weight and API-only, teams must accept complete dependency on Alibaba Cloud. Practitioners will also need to implement their own client-side reconnection logic, playback buffering, and voice-data governance policies to handle speaker consent.
This is our own summary of reporting by AlphaSignal



