Kannada Named Entity Recognition Using Deep Learning Techniques
Authors
Associate Professor, Department of Computer Science, Govt Women’s College Hunsuru, (India)
Associate Professor, Department of Computer Science, Sri Mahadeshwara Govt First grade Kollegal, Karnataka (India)
Associate Professor, Department of Computer Science, Govt Women’s College Hunsuru, (India)
Article Information
DOI: 10.51244/IJRSI.2026.1307000434
Subject Category: Computer Science
Volume/Issue: 13/7 | Page No: 5934-5941
Publication Timeline
Submitted: 2026-08-05
Accepted: 2026-08-10
Published: 2026-08-24
Abstract
Named Entity Recognition (NER) is a natural language processing task concerned with identifying mentions of named entities and classifying them according to a predefined set of categories. Despite the success of NER in domains, where such data is abundant it remains a formidable challenge for low-resource languages such as Kannada. In this paper we discuss the possible ways to approach NER for the Kannada language.
We explore various research directions including rule-based methods statistical machine learning neural networks and transformers based tagging methodologies. We highlight the various challenges in achieving NER for such a language and propose a transformer based contextual tagging framework for labelling sequences.
We propose to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task. We discuss various aspects for experimentation including data collection labelling data preparation methods data-splits evaluation metrics comparison with other models hyper parameter tuning entity-wise analysis and error analysis.
Keywords
Kannada, Named Entity Recognition, Natural Language Processing, Low-Resource Languages, Deep Learning, Transformer, BERT, IndicBERT, XLM-RoBERTa
Downloads
References
1. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT. [Google Scholar] [Crossref]
2. Wolf, T., et al. (2020). Transformers: State-of-the-Art Natural Language Processing. EMNLP. [Google Scholar] [Crossref]
3. Kakwani, D., et al. (2020). IndicNLPs: Benchmark and Resources for Indian Languages. [Google Scholar] [Crossref]
4. Lample, G., et al. (2016). Neural Architectures for Named Entity Recognition. NAACL. [Google Scholar] [Crossref]
5. Ma, X. & Hovy, E. (2016). End-to-end Sequence Labeling via Bi [Google Scholar] [Crossref]
6. Amarappa, S., & Sathyanarayana, S. V. (2015). Kannada Named Entity Recognition and Classification (NERC) Based on Multinomial Naïve Bayes (MNB) Classifier. The study reports experiments using a Kannada training corpus of 95,170 tokens and a test corpus of 5,000 tokens, with an F1-measure of 81%. [Google Scholar] [Crossref]
7. Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440–8451. [Google Scholar] [Crossref]
8. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171–4186. [Google Scholar] [Crossref]
9. Kakwani, D., Kunchukuttan, A., Golla, S., N. C., G., Bhattacharyya, A., Khapra, M. M., & Kumar, P. (2020). IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages. Findings of the Association for Computational Linguistics: EMNLP 2020. [Google Scholar] [Crossref]
10. Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K., & Dyer, C. (2016). Neural Architectures for Named Entity Recognition. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 260–270. [Google Scholar] [Crossref]
11. Ma, X., & Hovy, E. (2016). End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1064–1074. [Google Scholar] [Crossref]
12. Kunchukuttan, A., Kakwani, D., Golla, S., N. C., G., Bhattacharyya, A., Khapra, M. M., & Kumar, P. (2020). AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages. This work describes a large-scale Indic-language corpus and associated word embeddings. [Google Scholar] [Crossref]
13. Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., & Rush, A. M. (2020). Transformers: State-of-the-Art Natural Language Processing. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38–45. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- What the Desert Fathers Teach Data Scientists: Ancient Ascetic Principles for Ethical Machine-Learning Practice
- Comparative Analysis of Some Machine Learning Algorithms for the Classification of Ransomware
- Comparative Performance Analysis of Some Priority Queue Variants in Dijkstra’s Algorithm
- Transfer Learning in Detecting E-Assessment Malpractice from a Proctored Video Recordings.
- Dual-Modal Detection of Parkinson’s Disease: A Clinical Framework and Deep Learning Approach Using NeuroParkNet