Sentiment Analysis for Kannada–English Code-Mixed Social Media Text
Authors
Associate Professor, Department of Computer Science, Govt Women’s College, Hunsur, Mysuru, Karnataka, India (India)
Associate Professor, Department of Computer Science, Lal Bahadur Shastri Government First Grade College, RT Nagar Bengaluru, India (India)
Associate Professor, Department of Computer Science, Govt Women’s College, Hunsuru, Mysuru, Karnataka, India (India)
Associate Professor and Head Department of BCA, Govt College for Women’s Autonomous, Mandya, Karnataka, India (India)
Article Information
Publication Timeline
Submitted: 2026-08-06
Accepted: 2026-08-11
Published: 2026-08-19
Abstract
The exponential rise of social media has led to the generation of a large amount of informal text. Especially, code-mixed languages have gained substantial popularity among social media users. In India, the code-mixed language Kannada-English is widely used in social media platforms. This informal and non-standard language form brings forward significant challenges to natural language processing (NLP) tasks like sentiment analysis. This paper presents a detailed analysis of sentiment analysis in kannada-english code-mixed social media text. A manually annotated dataset is created and classified into positive, negative, and neutral sentiment labels. Various kinds of machine learning, deep learning, and transformer-based models are evaluated. The results suggest that the transformer-based models trained on Indian language corpora outperform others. Furthermore, this paper provides a comprehensive set of observations regarding sentiment analysis of Indian language code-mixed social media text.
Keywords
Kannada NLP, Code-Mixed Text, Sentiment Analysis, Low-Resource Languages, Transformer Models, Social Media Analytics
Downloads
References
1. K. Bali, M. Choudhury, and Y. Vyas, “Word-level language identification in code-mixed social media text,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) – Workshop on Computational Approaches to Code Switching, Doha, Qatar, 2014, pp. 50–61. [Google Scholar] [Crossref]
2. B. R. Chakravarthi, R. Priyadharshini, N. Jose, and E. Sherly, “DravidianCodeMix: Sentiment analysis and offensive language identification in Dravidian code-mixed text,” Language Resources and Evaluation, Springer, vol. 56, no. 2, pp. 453–478, 2022, doi: 10.1007/s10579-022-09583-7. [Google Scholar] [Crossref]
3. J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960, doi: 10.1177/001316446002000104. [Google Scholar] [Crossref]
4. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), Minneapolis, MN, USA, 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423. [Google Scholar] [Crossref]
5. B. Gambäck and A. Das, “On the state of the art of code-mixed sentiment analysis,” Computational Linguistics, vol. 42, no. 2, pp. 227–265, 2016. [Google Scholar] [Crossref]
6. B. Han, P. Cook, and T. Baldwin, “Automatically constructing a normalization dictionary for microblogs,” in Proceedings of the 2012 Conference on Empirical Methods in Natural Language Processing (EMNLP), Jeju Island, Korea, 2012, pp. 421–432. [Google Scholar] [Crossref]
7. S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997, doi: 10.1162/neco.1997.9.8.1735. [Google Scholar] [Crossref]
8. T. Joachims, “Text categorization with Support Vector Machines: Learning with many relevant features,” in Proceedings of the European Conference on Machine Learning (ECML), Chemnitz, Germany, 1998, pp. 137–142, doi: 10.1007/BFb0026683. [Google Scholar] [Crossref]
9. “IndicBERT: A multilingual ALBERT model for Indian languages,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 2020, pp. 5307–5315. [Google Scholar] [Crossref]
10. B. Liu, Sentiment Analysis and Opinion Mining, Morgan & Claypool Publishers, 2012. [Google Scholar] [Crossref]
11. C. Myers-Scotton, Social Motivations for Codeswitching: Evidence from Africa, Oxford University Press, 1993. [Google Scholar] [Crossref]
12. B. Pang, L. Lee, and S. Vaithyanathan, “Thumbs up? Sentiment classification using machine learning techniques,” in Proceedings of the 2002 Conference on Empirical Methods in Natural Language Processing (EMNLP), Philadelphia, PA, USA, 2002, pp. 79–86. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Humor, Critique, and Companionship: Audience Reception of an Unconventional Tiktok Marriage Proposal on Youtube
- Affective Visibility: Monetizing Care Work and Emotional Labor Among Full-Time Mothers on Douyin
- Generation Z Perception Towards Ai and Human Copywriting in Maggi Advertisements
- Media, Hate Speech, and National Cohesion in Nigeria: An Empirical Study of Media Discourse and Security Implications (2018–2025)
- Comparative Analysis of Brand Copywriting in the Generative AI Advertising Era