Indian AI company Sarvam AI has introduced Saaras V4, a new speech recognition model that supports all 22 official Indian languages as well as global English accents. The model is available through an API and includes real-time streaming transcription, batch processing, and speaker diarization. Sarvam AI lists real-time transcription pricing at 30 Indian rupees per hour.
Saaras V4 uses an encoder-decoder architecture, with the decoder built on Sarvam-3B, the company’s in-house 3 billion-parameter hybrid state space language model. It offers five output modes: transcription, verbatim records, code-mixed output, romanized transliteration, and translation. The company also added a keyword prompting feature designed to improve recognition accuracy for specific terms.
According to Sarvam AI, Saaras V4 posted the best results across multiple benchmarks and recorded a lower error rate than Deepgram Nova-3 and GPT-4o Transcribe on noisy audio datasets. Those figures, however, were reported by the company itself, and no independent replication results have been published so far. The report was cited by MarkTechPost.
Techub News reported that Indian AI company Sarvam AI has released Saaras V4, a new speech recognition model covering all 22 official Indian languages and global English accents.
The model is offered through an API. It supports real-time streaming transcription, batch processing, and speaker diarization, with real-time transcription priced at 30 Indian rupees per hour.
Architecture and output modes
Saaras V4 is built on an encoder-decoder architecture. Its decoder is based on Sarvam-3B, Sarvam AI’s in-house 3 billion-parameter hybrid state space language model.
The model provides five output modes: transcription, verbatim records, code-mixed output, romanized transliteration, and translation. It also adds a keyword prompting feature intended to improve accuracy for specific terminology.
Vendor-reported benchmark results
Sarvam AI said Saaras V4 achieved the best results across multiple benchmarks and posted a lower error rate than Deepgram Nova-3 and GPT-4o Transcribe on noisy audio datasets.
The benchmark data, however, was reported by the vendor itself, and no independent replication results have been published so far.
The item cited MarkTechPost as the source.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.