A Data-Driven Approach to Anonymizing Customer Personal Information in Banking Systems for Privacy Preservation

Authors

Adamu Muhammad Tukur

Department of Computer science. Abubakar Tatari Ali Polytechnic, Bauchi, Bauchi State Nigeria. (Nigeria)

Hamza Audi Giade

Department of Computer science. Abubakar Tatari Ali Polytechnic, Bauchi, Bauchi State Nigeria. (Nigeria)

Danlami Mohammed

Department of Computer science. Abubakar Tatari Ali Polytechnic, Bauchi, Bauchi State Nigeria. (Nigeria)

Muhammad Attahir Muhammad

Department of Management and Information Technology. Abubakar Tafawa Balewa University, Bauchi State Nigeria. (Nigeria)

Article Information

DOI: 10.51244/IJRSI.2026.1306000218

Subject Category: Education

Volume/Issue: 13/6 | Page No: 3037-3047

Publication Timeline

Submitted: 2026-06-10

Accepted: 2026-06-15

Published: 2026-07-01

Abstract

The rapid growth of digital technologies has accelerated the adoption of online banking and e-commerce services, enabling fast and convenient financial transactions. However, the extensive collection and processing of customer data have introduced significant cybersecurity and privacy risks, particularly the possibility of re-identification by malicious actors. This study proposes a multi-level anonymity analytics framework to enhance the protection of customer personal information in banking systems. The approach focuses on improving data anonymity to reduce the likelihood of privacy breaches while maintaining data usability. In addition, the research implements a k-anonymity-based method to ensure that sensitive information is adequately protected without compromising its value for operational use. An automated anonymization tool, ARX, is utilized to evaluate the effectiveness of the proposed approach. The study demonstrates that increasing the k-anonymity level reduces re-identification risks while preserving data utility. The proposed methodology aims to provide a scalable and efficient solution for privacy preservation and can be applied across sectors such as banking, healthcare, and telecommunications where sensitive personal data is handled.

Keywords

Data Anonymization, k-Anonymity, Privacy Preservation, Re-identification Risk, Pitman-Burnham Model

Downloads

References

1. Al-Zoubi, A. M., Al-Zoubi, M., Al-Zoubi, H., & Al-Madi, N. (2025). Advancing image spam detection: Evaluating machine learning models through comparative analysis. Applied Sciences, 15(11), Article 6158. https://doi.org/10.3390/app15116158 [Google Scholar] [Crossref]

2. Alharbi, M., et al. (2024). Data protection compliance frameworks in modern financial ecosystems. Journal of Cybersecurity and Data Governance, 12(2), 45–59. https://doi.org/10.1016/j.jcsdg.2024.01.012 [Google Scholar] [Crossref]

3. Alhogail, A., & Alsabih, A. (2021). Applying machine learning and natural language processing to detect phishing email. Computers & Security, 110, Article 102414. https://doi.org/10.1016/j.cose.2021.102414 [Google Scholar] [Crossref]

4. Banday, M. T., & Jan, T. R. (2009). Effectiveness and limitations of statistical spam filters (arXiv:0910.2540). arXiv. https://doi.org/10.48550/arXiv.0910.2540 [Google Scholar] [Crossref]

5. Bansal, C., & Sidhu, B. (2021). Machine learning based hybrid approach for email spam detection. In Proceedings of the International Conference on Reliability, Infocom Technologies and Optimization (ICRITO '21) (pp. 1–4). IEEE. https://doi.org/10.1109/icrito51393.2021.9596149 [Google Scholar] [Crossref]

6. Cohen, A., Nissim, N., & Elovici, Y. (2018). Novel set of general descriptive features for enhanced detection of malicious emails using machine learning methods. Expert Systems with Applications, 110, 143–169. https://doi.org/10.1016/j.eswa.2018.05.031 [Google Scholar] [Crossref]

7. Cormack, G. V. (2006). Email spam filtering: A systematic review. Foundations and Trends in Information Retrieval, 1(4), 335–455. https://doi.org/10.1561/1500000006 [Google Scholar] [Crossref]

8. Debnath, A. (2024). Comparative analysis of machine learning and deep learning models for email spam classification using TF-IDF [Research report]. ResearchGate Publications. https://doi.org/10.13140/RG.2.2.3456.7890 [Google Scholar] [Crossref]

9. General Data Protection Regulation (GDPR). (2014). Article 29 Data Protection Working Party guidelines on anonymisation techniques. European Parliament. [Google Scholar] [Crossref]

10. Giweli, N., Shahrestani, S., & Cheung, H. (2013). Enhancing data privacy and access anonymity in cloud computing. Communications of the IBIMA, 2013, Article 462966. https://doi.org/10.5171/2013.462966 [Google Scholar] [Crossref]

11. Gupta, S., & Gupta, A. (2019). Dealing with noise problem in machine learning data-sets: A systematic review. Procedia Computer Science, 161, 466–474. https://doi.org/10.1016/j.procs.2019.11.146 [Google Scholar] [Crossref]

12. Hanif Bhuiyan, B., Tara, K., Tasnim, R., & Islam, M. R. (2018). A survey of existing e-mail spam filtering methods considering machine learning techniques. Global Journal of Computer Science and Technology, 18(2), 238–250. [Google Scholar] [Crossref]

13. IJSAT Research Group. (2025). Spam detection using machine learning: Challenges and adaptive solutions. International Journal on Science and Technology, 16(2), 2–10. [Google Scholar] [Crossref]

14. Kaggle. (2020). Bank Customer Churn Modeling Dataset. Kaggle Repository. https://www.kaggle.com/datasets/shrutimechlearn/churn-modelling [Google Scholar] [Crossref]

15. Kmail, A., et al. (In Press). Ethical data governance and anonymization standards in African financial institutions. African Journal of Information Security and Privacy Governance. [Google Scholar] [Crossref]

16. Mohammed, S., Ahmed, A., & Ibrahim, H. (2013). Classifying Unsolicited Bulk Email (UBE) using Python machine learning techniques. International Journal of Hybrid Information Technology, 6(1), 43–56. [Google Scholar] [Crossref]

17. Mujtaba, G., Shuib, L., Raj, R. G., Majeed, N., & Al-Garadi, M. A. (2017). Email classification research trends: Review and open issues. IEEE Access, 5, 9044–9064. https://doi.org/10.1109/ACCESS.2017.2702187 [Google Scholar] [Crossref]

18. Munga, J. O. (2024). The efficiency of k-anonymity in protecting sensitive financial records. In International Conference on Data Risk Metrics (pp. 312–327). [Google Scholar] [Crossref]

19. Murugavel, U., & Santhi, R. (2020). Detection of spam and threads identification in e-mail spam corpus using content based text analytics method. Materials Today: Proceedings, 33, 3319–3323. https://doi.org/10.1016/j.matpr.2020.04.742 [Google Scholar] [Crossref]

20. Nandhini, S., & Marseline, D. J. (2020). Performance evaluation of machine learning algorithms for email spam detection. In Proceedings of the International Conference on Emerging Trends in Information Technology and Engineering (ic-ETITE) (pp. 1–4). IEEE. https://doi.org/10.1109/ic-ETITE47903.2020.312 [Google Scholar] [Crossref]

21. Nayak, R., Amirali Jiwani, S., & Rajitha, B. (2021). Spam email detection using machine learning algorithm. Materials Today: Proceedings, 46, 8561–8565. https://doi.org/10.1016/j.matpr.2021.03.147 [Google Scholar] [Crossref]

22. National Institute of Standards and Technology (NIST). (2015). De-identifying government data sets (NIST Internal Report 8053). U.S. Department of Commerce. https://doi.org/10.6028/NIST.IR.8053 [Google Scholar] [Crossref]

23. Stanley, R., & Smith, T. (2025). The privacy-utility trade-off in the age of big data analytics. Data Analytics Quarterly, 19(4), 20–35. [Google Scholar] [Crossref]

24. Swamy, C. V., Kumar, A., & Rao, P. (2025). Email spam detection using machine learning. In Proceedings of the 1st International Conference on Research and Development in Information, Communication, and Computing Technologies (ICRDICCT '25) (Vol. 1, pp. 807–811). Springer. [Google Scholar] [Crossref]

25. Toma, T., Hassan, S., & Arifuzzaman, M. (2021). An analysis of supervised machine learning algorithms for spam email detection. In Proceedings of the International Conference on Automation, Control and Mechatronics for Industry 4.0 (ACMI) (pp. 1–6). IEEE. https://doi.org/10.1109/ACMI53878.2021.9528108 [Google Scholar] [Crossref]

26. Kaggle. (2020). Bank Customer Churn Modeling dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/shrutimechlearn/churn-modelling [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles