Sarvam AI has released Saaras V4, a speech-to-text model covering all 22 scheduled Indian languages and global English. Built on an encoder-decoder architecture with a custom 3B hybrid state-space language model, Saaras V4 introduces five output modes and a keyterm prompting feature to bias recognition. Sarvam AI reports state-of-the-art accuracy across benchmarks, including a 16.03% WER on IndicContextEval, though independent validation is pending. The model is deployable via API starting at ₹30 per hour.
Indian language support
- ▪As of September 26, 2026, Deepgram Nova-3 covers 11 scheduled Indian languages and ElevenLabs Scribe v2 covers 14, compared to Saaras V4's 22
- ▪Sarvam AI released Saaras V4, a speech recognition model that covers all 22 scheduled Indian languages and English, including global English accents
Model deployment and integration
- ▪Sarvam AI designed Saaras V4 to use the same request shape as Saaras v3, enabling users to switch models with a one-line change
- ▪Sarvam AI made Saaras V4 deployable via its API on September 26, 2026, though the model's weights are not public and self-hosting documentation currently only covers Saaras v3
Model architecture and design
- ▪Saaras V4 is designed as an encoder-decoder system where an audio encoder converts waveforms into embeddings and a temporal-downsampling adapter shortens the sequence of embeddings
- ▪The decoder in Saaras V4 is Sarvam-3B, a 3-billion-parameter hybrid state-space language model trained from scratch in-house by Sarvam AI
Output modes and transcription options
- ▪Saaras V4 offers five selectable output modes from a single model: transcribe, verbatim, codemix, translit, and translate
- ▪The codemix mode of Saaras V4 outputs native script with English words left in English, while the translit mode returns the full utterance in Latin script
Keyterm prompting feature and performance
- ▪On the IndicContextEval benchmark, Saaras V4 reported a 16.03% Word Error Rate in the L5 keyword-prompting setting, which Sarvam AI states is the lowest score on the benchmark
- ▪Saaras V4 introduces a keyterm prompting feature that allows users to pass a JSON list of up to 50 terms of 64 characters each to bias speech recognition
Performance on English and noisy datasets
- ▪On seven English datasets, Saaras V4 achieved the lowest average Word Error Rate among the models benchmarked by Sarvam AI, including Deepgram Nova-3
- ▪On the Kathbath Noisy dataset, Sarvam AI reported that Saaras V4's error rate, measured with LLM-WER, is under half that of Deepgram Nova-3 and GPT-4o Transcribe
Language identification and benchmark verification
- ▪All benchmark results for Saaras V4 are vendor-reported by Sarvam AI, and independent reproduction of these figures has not been published as of September 26, 2026
- ▪On verified IndicVoices utterances, Saaras V4's language identification error is 2.9% across the top 10 Indian languages and 5.22% across all 22 scheduled Indian languages
Pricing and streaming capabilities
- ▪Sarvam AI lists Saaras V4 speech-to-text pricing at ₹30 per hour for real-time, streaming, and batch processing, and ₹45 per hour when speaker diarization is included
- ▪Saaras V4 supports real-time streaming via WebSocket with partial results and a vendor-claimed time to first token of below 150 milliseconds
Debatable claims
- ▪A single multilingual model is superior to specialized single-language models for regional speech recognition
- ▪Vendor-reported benchmarks are insufficient to justify enterprise adoption of new AI models
- ▪Sarvam AI should open-source the weights of its Saaras V4 model
Story comments
Loading comments…