Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics

Multimodal models stories

Sep 26, 2026

Sarvam AI releases Saaras V4 speech model covering 22 Indian languages

Sarvam AI launched Saaras V4, a speech-to-text model supporting all 22 scheduled Indian languages plus global English accents, representing a significant expansion in multilingual AI capabilities for the region.

Sep 26, 2026·2 sources
00
Sep 24, 2026

ElevenLabs releases Eleven v4 speech model with improved expression control and 90 languages

ElevenLabs launched Eleven v4 and v4 Turbo speech models, offering more accurate direction cues for laughter and whispers, consistent voice quality across long productions, and support for over 90 languages. The Turbo variant delivers 150-millisecond latency for real-time voice agents.

Sep 24, 2026·4 sources
00
Sep 13, 2026

Multimodal AI model outperforms established biomarkers for immunotherapy in lung cancer

A large international real-world study found that a multimodal explainable AI decision support tool outperformed established biomarkers for predicting immunotherapy outcomes in non-small cell lung cancer and improved physician decision-making.

Sep 13, 2026·2 sources
00
Sep 7, 2026

Alibaba releases Qwen-Drive 1.0 AI model with spatial awareness limitations

Alibaba's research arm has released Qwen-Drive 1.0, an AI model that integrates environmental perception, traffic Q&A, and route planning. The research reveals that text-image models don't automatically understand three-dimensional space and require specific training for spatial awareness.

Sep 7, 2026·1 source
00
Sep 6, 2026

H Company releases NeoMME multimodal encoders for visual document retrieval

H Company has launched NeoMME, a family of 260M and 800M parameter single-tower multimodal encoders designed for visual document retrieval. The models represent a departure from existing approaches that repurpose generative vision-language models as encoders.

Sep 6, 2026·1 source
00

H Company releases NeoMME multimodal encoders for visual document retrieval

H Company has launched NeoMME, a family of 260M and 800M parameter single-tower multimodal encoders designed for visual document retrieval. The models represent a departure from existing approaches that repurpose generative vision-language models as encoders.

Sep 6, 2026·1 source
00
Sep 1, 2026

Meta Superintelligence Labs releases Muse Voice Transcribe, a unified real-time transcription model

Meta Superintelligence Labs has released its first product, Muse Voice Transcribe, a real-time AI model that combines speech transcription, speaker diarization, and endpointing in a single system. The model can distinguish between multiple speakers and languages simultaneously, replacing the traditional approach of using three separate systems.

Sep 1, 2026·3 sources
00
Aug 26, 2026

Google releases Gemini 3.5 Transcribe speech-to-text model with 2.6% error rate across 85+ languages

Google has launched Gemini 3.5 Transcribe, a new AI-powered speech-to-text model that automatically removes filler words and cleans up speech in real-time. The model reports a 2.6% average word error rate across more than 85 languages and will be integrated into Chrome's voice typing feature.

Aug 26, 2026·4 sources
00

MiniMax revenue surges 283% in first half of 2026 on enterprise AI demand

Chinese AI firm MiniMax reported first-half 2026 revenue of $116.6 million, up 283% year-over-year, driven by a 700% jump in enterprise business. The Shanghai-based multimodal AI company's growth comes amid intense competition in China's AI sector.

Aug 26, 2026·3 sources
00
Aug 21, 2026

DeepSeek launches V4 Flash Vision model claiming performance near Anthropic's Opus 4.8

Chinese AI company DeepSeek released an experimental multimodal language model with visual comprehension capabilities, claiming its performance approaches that of Anthropic's advanced Opus 4.8 model. According to DeepSeek's published benchmarks, the new V4 Flash Vision model outperforms Opus 4.8 on three out of several tested metrics.

Aug 21, 2026·3 sources
00

DeepSeek launches vision-enabled AI model and cuts weekend API pricing

Chinese AI company DeepSeek released an experimental multimodal model claiming performance close to Anthropic's Opus-4.8, while simultaneously announcing it will charge off-peak API rates for all weekend usage starting August 23.

Aug 21, 2026·3 sources
00
Aug 15, 2026

New benchmark shows AI models struggle with visual perception, none reach 60% accuracy

Moonshot AI's PerceptionBench reveals that leading multimodal AI models, including GPT-5.6 Sol, perform poorly at basic visual perception tasks when separated from logical reasoning, with no frontier model achieving 60% accuracy. The benchmark demonstrates that many errors attributed to reasoning actually occur during the image-reading stage.

Aug 15, 2026·1 source
00
Aug 12, 2026

Google's AMIE medical AI matches primary care physicians in simulated video consultations

Google Research and DeepMind's AMIE system, a Gemini-based multi-agent AI, achieved performance on par with or exceeding primary care physicians in a randomized study of 100 simulated clinical video consultations with professional patient actors. The system combines real-time audio-visual perception, clinical reasoning, and low-latency dialogue capabilities.

Aug 12, 2026·3 sources
00
Aug 4, 2026

Mistral AI Releases Shieldstral, 3B-Parameter Open-Source Safety Classifier for AI Content Moderation

French AI startup Mistral AI announced Shieldstral on August 4, 2026, a 3-billion parameter open-weights multimodal safety classifier that matches models up to seven times its size in performance while allowing operators to define custom moderation policies in natural language at runtime.

Aug 4, 2026·9 sources
00
Jul 31, 2026

Chinese AI firm MiniMax launches H3 multimodal video generation model

MiniMax released its H3 video-generation model capable of processing text, images, video and audio, intensifying competition in China's AI video generation market currently led by ByteDance and Kuaishou.

Jul 31, 2026·1 source
00
Jul 26, 2026

Black Forest Labs releases FLUX 3 multimodal foundation model

Black Forest Labs launched FLUX 3, a multimodal foundation model that processes images, videos, and audio in a single architecture. It is the first FLUX model to support video, audio, and robot action prediction.

Jul 26, 2026·1 source
00
Jul 15, 2026

Thinking Machines Lab Releases Inkling, 975B-Parameter Open-Weights Multimodal AI Model

AI startup Thinking Machines Lab, led by former OpenAI CTO Mira Murati, has released Inkling, a 975-billion parameter mixture-of-experts model with open weights designed for enterprise customization and on-premises deployment. The company explicitly positions the model as focused on customizability rather than benchmark dominance.

Jul 15, 2026·7 sources
00
Jun 3, 2026

Google releases open-source Gemma 4 12B AI model capable of running locally on 16GB laptops

Google has launched Gemma 4 12B, an open-source AI model that can analyze audio and video while running entirely on enterprise laptops with 16GB of RAM, marking a focus on smaller, locally-deployable models rather than larger cloud-based alternatives.

Jun 3, 2026·5 sources
00
May 20, 2026

Alibaba Qwen Team Releases Real-Time Multimodal Translation Model Covering 60 Languages

Alibaba's Qwen team released Qwen3.5-LiveTranslate-Flash, a real-time multimodal translation model that processes audio and video simultaneously across 60 input languages with 2.8-second latency.

May 20, 2026·1 source
00
May 1, 2026

IBM Releases Granite 4.1 Model Family Across Language, Vision, Speech, and Embedding

IBM has launched its most expansive model release to date with the Granite 4.1 family, covering new language, vision, speech, embedding, and guardian models tailored for enterprise workloads.

May 1, 2026·2 sources
00
Apr 29, 2026

OpenMOSS Releases MOSS-Audio Open-Source Foundation Model

OpenMOSS released MOSS-Audio, an open-source foundation model that unifies speech, environmental sound, music, and temporal reasoning into a single architecture, outperforming existing open-source models on general audio benchmarks.

Apr 29, 2026·1 source
00
Apr 14, 2026

MiniMax Open Sources M2.7 Self-Evolving Agent Model Achieving 56.22% on SWE-Pro Benchmark

MiniMax released model weights for MiniMax M2.7 on Hugging Face, a self-evolving agent model that scores 56.22% on SWE-Pro and 57.0% on Terminal Bench 2. The company also released MMX-CLI, a command-line interface providing native access to image, video, speech, music, vision, and search capabilities.

Apr 14, 2026·2 sources
00
Apr 12, 2026

AI Models Prefer Guessing Over Asking for Help When Information Missing, Research Shows

ProactiveBench testing of 22 multimodal language models found that almost none ask users for help when visual information is missing, instead choosing to guess. Simple reinforcement learning can improve this behavior.

Apr 12, 2026·1 source
00
Apr 10, 2026

Google Gemini Launches Interactive 3D Model and Simulation Generation

Google upgraded Gemini to generate interactive 3D models and simulations in response to user questions, allowing users to rotate and manipulate AI-generated visualizations directly in the chat interface.

Apr 10, 2026·2 sources
00
Apr 3, 2026

Microsoft Launches Three New Foundation Models Including MAI-Transcribe-1

Microsoft's MAI division released three new foundation models: MAI-Transcribe-1 (speech-to-text in 25 languages, 2.5x faster than Azure Fast), MAI-Voice-1 (audio generation), and an image generation model. The transcription model costs $0.36 per audio hour.

Apr 3, 2026·3 sources
00

Top claims

  • ▪Vendor-reported benchmarks are insufficient to justify enterprise adoption of new AI models
  • ▪On seven English datasets, Saaras V4 achieved the lowest average Word Error Rate among the models benchmarked by Sarvam AI, including Deepgram Nova-3
  • ▪Sarvam AI designed Saaras V4 to use the same request shape as Saaras v3, enabling users to switch models with a one-line change

People involved

Mira Murati

Subtopics

AI foundation models8AI research & benchmarks8AI tools & products8AI startups7AI voice & speech tools6Open-source AI6AI agents4Business & enterprise AI3Generative audio3AI assistants & chatbots2AI safety benchmarks2China2China tech & industrial strategy2Document understanding2Open model families2Open weights vs closed models2Retrieval-augmented generation (RAG)23D generation1AI antitrust & competition1AI content watermarking & provenance1AI for developers1AI image generators1AI images & videos1AI in healthcare1

Related timelines

AI Data Center Gold Rush

101 stories

Congress

108 stories

Crypto hacks

100 stories

Ebola outbreak

58 stories

Iran War

209 stories