ElevenLabs launched Eleven v4 and v4 Turbo speech models, offering more accurate direction cues for laughter and whispers, consistent voice quality across long productions, and support for over 90 languages. The Turbo variant delivers 150-millisecond latency for real-time voice agents.
A large international real-world study found that a multimodal explainable AI decision support tool outperformed established biomarkers for predicting immunotherapy outcomes in non-small cell lung cancer and improved physician decision-making.
Alibaba's research arm has released Qwen-Drive 1.0, an AI model that integrates environmental perception, traffic Q&A, and route planning. The research reveals that text-image models don't automatically understand three-dimensional space and require specific training for spatial awareness.
H Company has launched NeoMME, a family of 260M and 800M parameter single-tower multimodal encoders designed for visual document retrieval. The models represent a departure from existing approaches that repurpose generative vision-language models as encoders.
H Company has launched NeoMME, a family of 260M and 800M parameter single-tower multimodal encoders designed for visual document retrieval. The models represent a departure from existing approaches that repurpose generative vision-language models as encoders.
Meta Superintelligence Labs has released its first product, Muse Voice Transcribe, a real-time AI model that combines speech transcription, speaker diarization, and endpointing in a single system. The model can distinguish between multiple speakers and languages simultaneously, replacing the traditional approach of using three separate systems.
Google has launched Gemini 3.5 Transcribe, a new AI-powered speech-to-text model that automatically removes filler words and cleans up speech in real-time. The model reports a 2.6% average word error rate across more than 85 languages and will be integrated into Chrome's voice typing feature.
Chinese AI firm MiniMax reported first-half 2026 revenue of $116.6 million, up 283% year-over-year, driven by a 700% jump in enterprise business. The Shanghai-based multimodal AI company's growth comes amid intense competition in China's AI sector.
Chinese AI company DeepSeek released an experimental multimodal language model with visual comprehension capabilities, claiming its performance approaches that of Anthropic's advanced Opus 4.8 model. According to DeepSeek's published benchmarks, the new V4 Flash Vision model outperforms Opus 4.8 on three out of several tested metrics.
Chinese AI company DeepSeek released an experimental multimodal model claiming performance close to Anthropic's Opus-4.8, while simultaneously announcing it will charge off-peak API rates for all weekend usage starting August 23.
Moonshot AI's PerceptionBench reveals that leading multimodal AI models, including GPT-5.6 Sol, perform poorly at basic visual perception tasks when separated from logical reasoning, with no frontier model achieving 60% accuracy. The benchmark demonstrates that many errors attributed to reasoning actually occur during the image-reading stage.
Google Research and DeepMind's AMIE system, a Gemini-based multi-agent AI, achieved performance on par with or exceeding primary care physicians in a randomized study of 100 simulated clinical video consultations with professional patient actors. The system combines real-time audio-visual perception, clinical reasoning, and low-latency dialogue capabilities.
French AI startup Mistral AI announced Shieldstral on August 4, 2026, a 3-billion parameter open-weights multimodal safety classifier that matches models up to seven times its size in performance while allowing operators to define custom moderation policies in natural language at runtime.
AI startup Thinking Machines Lab, led by former OpenAI CTO Mira Murati, has released Inkling, a 975-billion parameter mixture-of-experts model with open weights designed for enterprise customization and on-premises deployment. The company explicitly positions the model as focused on customizability rather than benchmark dominance.
Google has launched Gemma 4 12B, an open-source AI model that can analyze audio and video while running entirely on enterprise laptops with 16GB of RAM, marking a focus on smaller, locally-deployable models rather than larger cloud-based alternatives.
MiniMax released model weights for MiniMax M2.7 on Hugging Face, a self-evolving agent model that scores 56.22% on SWE-Pro and 57.0% on Terminal Bench 2. The company also released MMX-CLI, a command-line interface providing native access to image, video, speech, music, vision, and search capabilities.
ProactiveBench testing of 22 multimodal language models found that almost none ask users for help when visual information is missing, instead choosing to guess. Simple reinforcement learning can improve this behavior.
Microsoft's MAI division released three new foundation models: MAI-Transcribe-1 (speech-to-text in 25 languages, 2.5x faster than Azure Fast), MAI-Voice-1 (audio generation), and an image generation model. The transcription model costs $0.36 per audio hour.