IIIT Hyderabad Publications |
|||||||||
|
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji PredictionAuthors: Nurendra Choudhary,Rajat Singh,Vijjini Anvesh Rao,Manish Shrivastava Conference: 27th International Conference on Computational Linguistics (COLING 2018) (COLING-2018 2018) Location Santa Fe, New Mexico, USA Date: 2018-08-20 Report no: IIIT/TR/2018/136 AbstractIn this paper, we leverage social media platforms such as twitter for developing corpus across multiple languages. The corpus creation methodology is applicable for resource-scarce languages provided the speakers of that particular language are active users on social media platforms. We present an approach to extract social media microblogs such as tweets (Twitter). In this paper, we create corpus for multilingual sentiment analysis and emoji prediction in Hindi, Bengali and Telugu. Further, we perform and analyze multiple NLP tasks utilizing the corpus to get interesting observations. Full paper: pdf Centre for Language Technologies Research Centre |
||||||||
Copyright © 2009 - IIIT Hyderabad. All Rights Reserved. |