620063F2A720330A989DF157F6AC4DDC

IISc researchers release open-source AI model supporting 65 Indian languages, dialect

Bengaluru, Aug 14 (INB) Researchers at the Indian Institute of Science’s (IISc) SPIRE Lab, in collaboration with ARTPARK and with support from Google, have developed and released an open-source speech recognition model capable of supporting 65 Indian languages and dialects, an official statement said. This include several languages and dialects underserved by existing speech AI systems, the IISc said in a statement on Thursday. Named SraVaani, the model is described as the first multilingual Indian speech recognition model trained on 65 languages and dialects, including more than 40 regional languages and dialects that are not officially supported by most existing speech recognition systems, according to IISc. The model covers 20 scheduled languages and 45 regional languages and dialects, and is aimed at expanding access to speech-to-text technology to potentially around 25 crore people whose languages are not adequately supported by existing systems, the institute said. SraVaani supports languages including Garo, Angika, Chakma, Kokborok, Tulu, Bundeli and Bajjika, besides English and Sanskrit. It can generate text in 10 scripts and automatically identify the language being spoken, eliminating the need for users to specify a language tag in advance, according to the statement. The model has been made freely available on Hugging Face under an MIT licence, along with a demonstration and fine-tuning code, allowing developers and researchers to adapt it for applications involving regional-language speech recognition, low-resource languages and sovereign AI, IISc said. “SraVaani, serving more than 60 Indian languages, is a contribution” towards India’s inclusive AI ambitions, IISc Director Prof Govindan Rangarajan said. Prof Prasanta Kumar Ghosh, Professor at IISc and Principal Investigator of Project Vaani, said the objective behind the initiative was to ensure that voice AI works for Indians irrespective of the resources available for their languages. The model was evaluated across eight public benchmark datasets and achieved accuracy comparable to leading Indic speech recognition systems for widely supported Indian languages, the statement said. Its performance was particularly notable for underserved languages. For instance, SraVaani recorded a word error rate of 9.5 per cent for Garo, compared with 69.4 per cent for the next-best system evaluated, according to IISc. The model’s foundation is the Vaani dataset, developed under Project Vaani. Researchers have collected more than 31,000 hours of spontaneous speech from 1.56 lakh people across 165 districts in 28 states and three Union Territories. Speakers were asked to describe images in their own words rather than read prepared sentences, enabling the dataset to capture natural speech, dialects and regional variations, IISc said. SraVaani was trained in three stages using the Vaani corpus, 11.8 million audio-image pairs and transcribed speech from Vaani and other publicly available Indian speech datasets.

Leave a Reply

Your email address will not be published. Required fields are marked *