
Researchers at the Indian Institute of Science (IISc) have released SraVaani, an open-source voice AI model that recognises speech in 65 Indian languages and dialects. The model converts spoken words into text…
Researchers at the Indian Institute of Science (IISc) have released SraVaani, an open-source voice AI model that recognises speech in 65 Indian languages and dialects. The model converts spoken words into text across 10 scripts and automatically identifies the language being used, without requiring manual selection. SraVaani covers 20 scheduled languages and 45 regional languages, including Garo, Angika, and Tulu, which are not supported by existing speech-recognition systems. Together, these languages are spoken by about 25 crore people as per the 2011 census.
The model is freely available on Hugging Face under an MIT licence. IISc said SraVaani performed on par with leading Indian speech systems on commonly supported languages and showed a clear advantage on underserved ones. On Garo, its word error rate was 9.5 per cent against 69.4 per cent for the next-best system. SraVaani was trained on more than 31,000 hours of speech collected from 156,000 people across 165 districts in 28 states under Project Vaani.
SraVaani is a welcome step, but the usual hype around 'AI for all' should be watched. The claim of covering 65 languages sounds impressive, yet the data is nearly a decade old and many of those languages have tiny speaker bases. The real test is whether startups and government services actually deploy this model. Will a Garo speaker see a bank app that understands her? That is the only number that will settle the question.
Source: timesofindia.indiatimes.com
This story was synthesised by AI from the source linked above.