Alibaba's Qwen team released Qwen3.5-LiveTranslate-Flash on May 20, 2026, a real-time multimodal translation model covering 60 input languages with speech output in 29 languages. The model processes audio and video simultaneously at 2.8-second latency and adds real-time speaker voice cloning, vision-enhanced comprehension via lip movements and on-screen text, and dynamic keyword configuration over the previous Qwen3 version.
Qwen3.5-LiveTranslate-Flash release
- ▪Alibaba's Qwen team released Qwen3.5-LiveTranslate-Flash on May 20, 2026
- ▪Qwen3.5-LiveTranslate-Flash is a real-time multimodal translation model
- ▪Qwen3.5-LiveTranslate-Flash covers 60 input languages
- ▪Qwen3.5-LiveTranslate-Flash produces speech output in 29 languages
Multimodal translation capabilities
- ▪Qwen3.5-LiveTranslate-Flash operates at 2.8 seconds of latency
- ▪Qwen3.5-LiveTranslate-Flash includes vision-enhanced comprehension via lip movements and on-screen text as a key addition over the previous Qwen3 version
- ▪Qwen3.5-LiveTranslate-Flash includes dynamic keyword configuration for domain-specific terminology as a key addition over the previous Qwen3 version
- ▪Qwen3.5-LiveTranslate-Flash processes audio and video simultaneously
- ▪Qwen3.5-LiveTranslate-Flash includes real-time speaker voice cloning as a key addition over the previous Qwen3 version
Performance benchmarks
- ▪Qwen3.5-LiveTranslate-Flash outperforms major commercial alternatives on FLEURS benchmark
API availability
- ▪Qwen3.5-LiveTranslate-Flash uses a WebSocket-based protocol for API access
- ▪Qwen3.5-LiveTranslate-Flash is available as an API-only model through Alibaba Cloud Model Studio
Story comments
Loading comments…