Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Alibaba Qwen Team Releases Real-Time Multimodal Translation Model Covering 60 Languages
00

Alibaba Qwen Team Releases Real-Time Multimodal Translation Model Covering 60 Languages

May 20, 2026

Alibaba's Qwen team released Qwen3.5-LiveTranslate-Flash on May 20, 2026, a real-time multimodal translation model covering 60 input languages with speech output in 29 languages. The model processes audio and video simultaneously at 2.8-second latency and adds real-time speaker voice cloning, vision-enhanced comprehension via lip movements and on-screen text, and dynamic keyword configuration over the previous Qwen3 version.

Qwen3.5-LiveTranslate-Flash release

  • ▪Alibaba's Qwen team released Qwen3.5-LiveTranslate-Flash on May 20, 2026
  • ▪Qwen3.5-LiveTranslate-Flash is a real-time multimodal translation model
  • ▪Qwen3.5-LiveTranslate-Flash covers 60 input languages
  • ▪Qwen3.5-LiveTranslate-Flash produces speech output in 29 languages

Multimodal translation capabilities

  • ▪Qwen3.5-LiveTranslate-Flash operates at 2.8 seconds of latency
  • ▪Qwen3.5-LiveTranslate-Flash includes vision-enhanced comprehension via lip movements and on-screen text as a key addition over the previous Qwen3 version
  • ▪Qwen3.5-LiveTranslate-Flash includes dynamic keyword configuration for domain-specific terminology as a key addition over the previous Qwen3 version
  • ▪Qwen3.5-LiveTranslate-Flash processes audio and video simultaneously
  • ▪Qwen3.5-LiveTranslate-Flash includes real-time speaker voice cloning as a key addition over the previous Qwen3 version

Performance benchmarks

  • ▪Qwen3.5-LiveTranslate-Flash outperforms major commercial alternatives on FLEURS benchmark

API availability

  • ▪Qwen3.5-LiveTranslate-Flash uses a WebSocket-based protocol for API access
  • ▪Qwen3.5-LiveTranslate-Flash is available as an API-only model through Alibaba Cloud Model Studio

1 source

Marktechpost
Alibaba Qwen Team Introduces Qwen3.5-LiveTranslate-Flash: Real-Time Multimodal Interpretation Across 60 Languages at 2.8-Second Latency
View source article

Featured stories

View more in Large language models (LLMs)

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

Anthropic releases Claude Sonnet 5.5 with 30% speed and cost improvements ahead of planned IPO

Sep 28, 2026 · 6 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

Story comments

Loading comments…

Related Projects

Qwen

Topics

Large language models (LLMs)AI voice & speech toolsAI tools & productsMultimodal models

Featured stories

View more in Large language models (LLMs)

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

Anthropic releases Claude Sonnet 5.5 with 30% speed and cost improvements ahead of planned IPO

Sep 28, 2026 · 6 sources

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sep 26, 2026 · 2 sources

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources