Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition

Authors: Ayushi Pandey,Srivastava B M L,Rohit Kumar,Bhanu Teja Nellore,Teja K S,Suryakanth V Gangashetty
Conference: 11th edition of the Language Resources and Evaluation Conference (LREC-2018 2018)
Location Miyazaki (Japan)
Date: 2018-05-07
Report no: IIIT/TR/2018/13

Abstract

The paper presents the development of a phonetically balanced read speech corpus of code-mixed Hindi-English. Phonetic balance in the corpus has been created by selecting sentences that contained triphones lower in frequency than a predefined threshold. The assumption with a compulsory inclusion of such rare units was that the high frequency triphones will inevitably be included. Using this metric, the Pearson’s correlation coefficient of the phonetically balanced corpus with a large code-mixed reference corpus was recorded to be 0.996. The data for corpus creation has been extracted from selected sections of Hindi newspapers.These sections contain frequent English insertions in a matrix of Hindi sentence. Statistics on the phone and triphone distribution have been presented, to graphically display the phonetic likeness between the reference corpus and the corpus sampled through our method.

Full paper: pdf

Centre for Language Technologies Research Centre

IIIT Hyderabad Publications

Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition

Abstract