IIIT Hyderabad Publications |
|||||||||
|
Unsupervised Morphological Expansion of Small Datasets for Improving Word EmbeddingsAuthors: Arihant Gupta,Sarfaraz Akhtar Syed,Avijit Vajpayee,Arjit Srivastava,Manish Shrivastava Conference: 18th International Conference on Computational Linguistics and Intelligent Text Processing (CICling-2017 2017) Location Budapest, Hungary Date: 2017-04-17 Report no: IIIT/TR/2017/30 AbstractWe present a language independent, unsupervised method for building word embeddings using morphological expansion of text. Our model handles the problem of data sparsity and yields improved word embeddings by relying on training word embeddings on artificially generated sentences. We evaluate our method using small sized training sets on eleven test sets for the word similarity task across seven languages. Further, for English, we evaluated the impacts of our approach using a large training set on three standard test sets. Our method improved results across all languages. Full paper: pdf Centre for Language Technologies Research Centre |
||||||||
Copyright © 2009 - IIIT Hyderabad. All Rights Reserved. |