Kannada Named Entity Recognition Using Deep Learning Techniques

Authors

Dr. Pushpalatha M

Associate Professor, Department of Computer Science, Govt Women’s College Hunsuru, (India)

Dr. Hemakumar G

Associate Professor, Department of Computer Science, Sri Mahadeshwara Govt First grade Kollegal, Karnataka (India)

Dr. Santhosh Kuamr B N

Associate Professor, Department of Computer Science, Govt Women’s College Hunsuru, (India)

Article Information

DOI: 10.51244/IJRSI.2026.1307000434

Subject Category: Computer Science

Volume/Issue: 13/7 | Page No: 5934-5941

Publication Timeline

Submitted: 2026-08-05

Accepted: 2026-08-10

Published: 2026-08-24

Abstract

Named Entity Recognition (NER) is a natural language processing task concerned with identifying mentions of named entities and classifying them according to a predefined set of categories. Despite the success of NER in domains, where such data is abundant it remains a formidable challenge for low-resource languages such as Kannada. In this paper we discuss the possible ways to approach NER for the Kannada language.
We explore various research directions including rule-based methods statistical machine learning neural networks and transformers based tagging methodologies. We highlight the various challenges in achieving NER for such a language and propose a transformer based contextual tagging framework for labelling sequences.
We propose to use mBERT IndicBERT and XLM-RoBERTa language models pretrained on target and other related Indic language corpora and further fine-tune these models for the NER task. We discuss various aspects for experimentation including data collection labelling data preparation methods data-splits evaluation metrics comparison with other models hyper parameter tuning entity-wise analysis and error analysis.

Keywords

Kannada, Named Entity Recognition, Natural Language Processing, Low-Resource Languages, Deep Learning, Transformer, BERT, IndicBERT, XLM-RoBERTa

Downloads

References

1. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT. [Google Scholar] [Crossref]

2. Wolf, T., et al. (2020). Transformers: State-of-the-Art Natural Language Processing. EMNLP. [Google Scholar] [Crossref]

3. Kakwani, D., et al. (2020). IndicNLPs: Benchmark and Resources for Indian Languages. [Google Scholar] [Crossref]

4. Lample, G., et al. (2016). Neural Architectures for Named Entity Recognition. NAACL. [Google Scholar] [Crossref]

5. Ma, X. & Hovy, E. (2016). End-to-end Sequence Labeling via Bi [Google Scholar] [Crossref]

6. Amarappa, S., & Sathyanarayana, S. V. (2015). Kannada Named Entity Recognition and Classification (NERC) Based on Multinomial Naïve Bayes (MNB) Classifier. The study reports experiments using a Kannada training corpus of 95,170 tokens and a test corpus of 5,000 tokens, with an F1-measure of 81%. [Google Scholar] [Crossref]

7. Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440–8451. [Google Scholar] [Crossref]

8. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171–4186. [Google Scholar] [Crossref]

9. Kakwani, D., Kunchukuttan, A., Golla, S., N. C., G., Bhattacharyya, A., Khapra, M. M., & Kumar, P. (2020). IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages. Findings of the Association for Computational Linguistics: EMNLP 2020. [Google Scholar] [Crossref]

10. Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K., & Dyer, C. (2016). Neural Architectures for Named Entity Recognition. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 260–270. [Google Scholar] [Crossref]

11. Ma, X., & Hovy, E. (2016). End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1064–1074. [Google Scholar] [Crossref]

12. Kunchukuttan, A., Kakwani, D., Golla, S., N. C., G., Bhattacharyya, A., Khapra, M. M., & Kumar, P. (2020). AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages. This work describes a large-scale Indic-language corpus and associated word embeddings. [Google Scholar] [Crossref]

13. Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., & Rush, A. M. (2020). Transformers: State-of-the-Art Natural Language Processing. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38–45. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles