Gender Disparities in AI-Driven Depression Detection: A Systematic Review of Algorithmic Bias and Its Implications for Women’s Health
Authors
Department of Public Health, University of Hertfordshire (UK)
Department of Clinical Psychology, University of South Wales (UK)
Department of Clinical Psychology, University of South Wales (UK)
Department of Public Health, National Open University of Nigeria (UK)
Department of Public Health, University of Hertfordshire (Nigeria)
Article Information
DOI: 10.47772/IJRISS.2026.100400058
Subject Category: Public Health
Volume/Issue: 10/4 | Page No: 813-820
Publication Timeline
Submitted: 2026-04-08
Accepted: 2026-04-13
Published: 2026-04-29
Abstract
Background: This systematic review critically synthesises evidence regarding gender disparities in AI-driven depression detection, emphasising algorithmic bias and its ramifications for women's health.
Methods: According to PRISMA guidelines, a systematic search of six databases (PubMed, IEEE Xplore, ACM Digital Library, PsycINFO, Scopus, and Web of Science) found 28 studies that met the criteria and were published between 2015 and 2025
Results: The results show that there is a significant disparity in performance between men and women, with models often being less sensitive to depression in women. Bias sources include underrepresentation of female subjects in training data, reliance on male-normative symptom presentation, and feature selection that neglects psychosocial determinants of women's mental health.
Conclusion: The review concludes that current AI models risk perpetuating diagnostic inequities, necessitating the development of gender-inclusive datasets and fairness-aware algorithms.
Keywords
Gender bias; Artificial intelligence
Downloads
References
1. Albert, P. (2015). Why is depression more prevalent in women? Journal of Psychiatry & Neuroscience, 40(4), 219–221. https://doi.org/10.1503/jpn.150205 [Google Scholar] [Crossref]
2. Calvo, R. A., Milne, D. N., Hussain, M. S., & Christensen, H. (2017). Natural language processing in mental health applications using non-clinical texts. Natural Language Engineering, 23(5), 649–685. https://doi.org/10.1017/s1351324916000383 [Google Scholar] [Crossref]
3. Chen, J.-M., Rao, M., Wei, Y.-T., Zhou, Q.-G., Tao, J.-L., Wang, S.-B., & Bi, B. (2025). Machine learning-based nomogram for predicting depressive symptoms in women: A cross-sectional study in Guangdong Province, China. World Journal of Psychiatry, 15(8). https://doi.org/10.5498/wjp.v15.i8.106622 [Google Scholar] [Crossref]
4. Diaz Ochoa, J. G., Layer, N., Mahr, J., Mustafa, F. E., Menzel, C. U., Müller, M., Schilling, T., Illerhaus, G., Knott, M., & Krohn, A. (2025). Optimized BERT-based NLP outperforms zero-shot methods for automated symptom detection in clinical practice. Frontiers in Digital Health, 7. https://doi.org/10.3389/fdgth.2025.1623922 [Google Scholar] [Crossref]
5. Harding, S. (1991). Whose Science? Whose knowledge? Cornell University Press. https://www.cornellpress.cornell.edu/book/9780801497469/whose-science-whose-knowledge/#bookTabs=1 [Google Scholar] [Crossref]
6. Hyde, J. S., & Mezulis, A. H. (2020). Gender differences in depression: Biological, affective, cognitive, and sociocultural factors. Harvard Review of Psychiatry, 28(1), 4–13. https://doi.org/10.1097/hrp.0000000000000230 [Google Scholar] [Crossref]
7. Kautzky, A., Dold, M., Bartova, L., Spies, M., Kranz, G. S., Souery, D., Montgomery, S., Mendlewicz, J., Zohar, J., Fabbri, C., Serretti, A., Lanzenberger, R., Dikeos, D., Rujescu, D., & Kasper, S. (2019). Clinical predictors of treatment resistant depression: Replication results from the European multicenter study. European Neuropsychopharmacology, 29, S61–S62. https://doi.org/10.1016/j.euroneuro.2018.11.1038 [Google Scholar] [Crossref]
8. Marcus, S. M., Young, E. A., Kerber, K. B., Kornstein, S., Farabaugh, A. H., Mitchell, J., Wisniewski, S. R., Balasubramani, G. K., Trivedi, M. H., & Rush, A. J. (2005). Gender differences in depression: Findings from the STAR*D study. Journal of Affective Disorders, 87(2-3), 141–150. https://doi.org/10.1016/j.jad.2004.09.008 [Google Scholar] [Crossref]
9. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607 [Google Scholar] [Crossref]
10. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342 [Google Scholar] [Crossref]
11. Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., & Moher, D. (2021). Updating guidance for reporting systematic reviews: development of the PRISMA 2020 statement. Journal of Clinical Epidemiology, 134(134), 103–112. https://doi.org/10.1016/j.jclinepi.2021.02.003 [Google Scholar] [Crossref]
12. Raj, A., Ali, Z., Chaudhary, S., Bali, K. K., & Sharma, A. (2024). Depression Detection Using BERT on Social Media Platforms. 2022 IEEE International Conference on Artificial Intelligence in Engineering and Technology (IICAIET), 228–233. https://doi.org/10.1109/iicaiet62352.2024.10730329 [Google Scholar] [Crossref]
13. Schnack, H. G., & Kahn, R. S. (2016). Detecting Neuroimaging Biomarkers for Psychiatric Disorders: Sample Size Matters. Frontiers in Psychiatry, 7. https://doi.org/10.3389/fpsyt.2016.00050 [Google Scholar] [Crossref]
14. Shatte, A. B. R., Hutchinson, D. M., & Teague, S. J. (2019). Machine learning in mental health: a scoping review of methods and applications. Psychological Medicine, 49(09), 1426–1448. https://doi.org/10.1017/s0033291719000151 [Google Scholar] [Crossref]
15. Thakkar, A., Gupta, A., & De Sousa, A. (2024). Artificial intelligence in positive mental health: a narrative review. Frontiers in Digital Health, 6(1280235). https://doi.org/10.3389/fdgth.2024.1280235 [Google Scholar] [Crossref]
16. Wolff, R. F., Moons, K. G. M., Riley, R. D., Whiting, P. F., Westwood, M., Collins, G. S., Reitsma, J. B., Kleijnen, J., & Mallett, S. (2019). PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Annals of Internal Medicine, 170(1), 51. https://doi.org/10.7326/m18-1376 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Tribal Child Nutrition and Health in District of Sundargarh: A Public Health Review of ICDS Intervention
- Knowledge, Attitudes and Practices Towards Prostate Cancer Screening Amongst Men Aged 40-60 Years in The Buea Health District: A Cross-Sectional Study
- Compliance with JCI Protocols: A Focus on Employee Safety
- Influence and Involvement of Teachers in Menstrual Hygiene Management of Female Secondary School Students in Kogi State, Nigeria
- A Critical Evaluation of Ayushman Bharat Pradhan Mantri Jan Arogya Yojana in Bihar