Sentiment Analysis for Kannada–English Code-Mixed Social Media Text

Authors

Dr. Pushpalatha M

Associate Professor, Department of Computer Science, Govt Women’s College, Hunsur, Mysuru, Karnataka, India (India)

Dr. Sasikala P

Associate Professor, Department of Computer Science, Lal Bahadur Shastri Government First Grade College, RT Nagar Bengaluru, India (India)

Dr. Santhosh Kuamr B N

Associate Professor, Department of Computer Science, Govt Women’s College, Hunsuru, Mysuru, Karnataka, India (India)

Dr. H S Nagalakshmi

Associate Professor and Head Department of BCA, Govt College for Women’s Autonomous, Mandya, Karnataka, India (India)

Article Information

DOI: 10.51244/IJRSI.2026.1307000354

Subject Category: Media

Volume/Issue: 13/7 | Page No: 4785-4790

Publication Timeline

Submitted: 2026-08-06

Accepted: 2026-08-11

Published: 2026-08-19

Abstract

The exponential rise of social media has led to the generation of a large amount of informal text. Especially, code-mixed languages have gained substantial popularity among social media users. In India, the code-mixed language Kannada-English is widely used in social media platforms. This informal and non-standard language form brings forward significant challenges to natural language processing (NLP) tasks like sentiment analysis. This paper presents a detailed analysis of sentiment analysis in kannada-english code-mixed social media text. A manually annotated dataset is created and classified into positive, negative, and neutral sentiment labels. Various kinds of machine learning, deep learning, and transformer-based models are evaluated. The results suggest that the transformer-based models trained on Indian language corpora outperform others. Furthermore, this paper provides a comprehensive set of observations regarding sentiment analysis of Indian language code-mixed social media text.

Keywords

Kannada NLP, Code-Mixed Text, Sentiment Analysis, Low-Resource Languages, Transformer Models, Social Media Analytics

Downloads

References

1. K. Bali, M. Choudhury, and Y. Vyas, “Word-level language identification in code-mixed social media text,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) – Workshop on Computational Approaches to Code Switching, Doha, Qatar, 2014, pp. 50–61. [Google Scholar] [Crossref]

2. B. R. Chakravarthi, R. Priyadharshini, N. Jose, and E. Sherly, “DravidianCodeMix: Sentiment analysis and offensive language identification in Dravidian code-mixed text,” Language Resources and Evaluation, Springer, vol. 56, no. 2, pp. 453–478, 2022, doi: 10.1007/s10579-022-09583-7. [Google Scholar] [Crossref]

3. J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960, doi: 10.1177/001316446002000104. [Google Scholar] [Crossref]

4. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), Minneapolis, MN, USA, 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423. [Google Scholar] [Crossref]

5. B. Gambäck and A. Das, “On the state of the art of code-mixed sentiment analysis,” Computational Linguistics, vol. 42, no. 2, pp. 227–265, 2016. [Google Scholar] [Crossref]

6. B. Han, P. Cook, and T. Baldwin, “Automatically constructing a normalization dictionary for microblogs,” in Proceedings of the 2012 Conference on Empirical Methods in Natural Language Processing (EMNLP), Jeju Island, Korea, 2012, pp. 421–432. [Google Scholar] [Crossref]

7. S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997, doi: 10.1162/neco.1997.9.8.1735. [Google Scholar] [Crossref]

8. T. Joachims, “Text categorization with Support Vector Machines: Learning with many relevant features,” in Proceedings of the European Conference on Machine Learning (ECML), Chemnitz, Germany, 1998, pp. 137–142, doi: 10.1007/BFb0026683. [Google Scholar] [Crossref]

9. “IndicBERT: A multilingual ALBERT model for Indian languages,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 2020, pp. 5307–5315. [Google Scholar] [Crossref]

10. B. Liu, Sentiment Analysis and Opinion Mining, Morgan & Claypool Publishers, 2012. [Google Scholar] [Crossref]

11. C. Myers-Scotton, Social Motivations for Codeswitching: Evidence from Africa, Oxford University Press, 1993. [Google Scholar] [Crossref]

12. B. Pang, L. Lee, and S. Vaithyanathan, “Thumbs up? Sentiment classification using machine learning techniques,” in Proceedings of the 2002 Conference on Empirical Methods in Natural Language Processing (EMNLP), Philadelphia, PA, USA, 2002, pp. 79–86. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles