Gender Disparities in AI-Driven Depression Detection: A Systematic Review of Algorithmic Bias and Its Implications for Women’s Health

Authors

Ihuoma Goodness Dike

Department of Public Health, University of Hertfordshire (UK)

Ndorenyin Saviour Udofia

Department of Clinical Psychology, University of South Wales (UK)

Dumebi Okuagu

Department of Clinical Psychology, University of South Wales (UK)

Anthony, Clement Ogbeh

Department of Public Health, National Open University of Nigeria (UK)

Sandra Ada Collins

Department of Public Health, University of Hertfordshire (Nigeria)

Article Information

DOI: 10.47772/IJRISS.2026.100400058

Subject Category: Public Health

Volume/Issue: 10/4 | Page No: 813-820

Publication Timeline

Submitted: 2026-04-08

Accepted: 2026-04-13

Published: 2026-04-29

Abstract

Background: This systematic review critically synthesises evidence regarding gender disparities in AI-driven depression detection, emphasising algorithmic bias and its ramifications for women's health.
Methods: According to PRISMA guidelines, a systematic search of six databases (PubMed, IEEE Xplore, ACM Digital Library, PsycINFO, Scopus, and Web of Science) found 28 studies that met the criteria and were published between 2015 and 2025
Results: The results show that there is a significant disparity in performance between men and women, with models often being less sensitive to depression in women. Bias sources include underrepresentation of female subjects in training data, reliance on male-normative symptom presentation, and feature selection that neglects psychosocial determinants of women's mental health.
Conclusion: The review concludes that current AI models risk perpetuating diagnostic inequities, necessitating the development of gender-inclusive datasets and fairness-aware algorithms.

Keywords

Gender bias; Artificial intelligence

Downloads

References

1. Albert, P. (2015). Why is depression more prevalent in women? Journal of Psychiatry & Neuroscience, 40(4), 219–221. https://doi.org/10.1503/jpn.150205 [Google Scholar] [Crossref]

2. Calvo, R. A., Milne, D. N., Hussain, M. S., & Christensen, H. (2017). Natural language processing in mental health applications using non-clinical texts. Natural Language Engineering, 23(5), 649–685. https://doi.org/10.1017/s1351324916000383 [Google Scholar] [Crossref]

3. Chen, J.-M., Rao, M., Wei, Y.-T., Zhou, Q.-G., Tao, J.-L., Wang, S.-B., & Bi, B. (2025). Machine learning-based nomogram for predicting depressive symptoms in women: A cross-sectional study in Guangdong Province, China. World Journal of Psychiatry, 15(8). https://doi.org/10.5498/wjp.v15.i8.106622 [Google Scholar] [Crossref]

4. Diaz Ochoa, J. G., Layer, N., Mahr, J., Mustafa, F. E., Menzel, C. U., Müller, M., Schilling, T., Illerhaus, G., Knott, M., & Krohn, A. (2025). Optimized BERT-based NLP outperforms zero-shot methods for automated symptom detection in clinical practice. Frontiers in Digital Health, 7. https://doi.org/10.3389/fdgth.2025.1623922 [Google Scholar] [Crossref]

5. Harding, S. (1991). Whose Science? Whose knowledge? Cornell University Press. https://www.cornellpress.cornell.edu/book/9780801497469/whose-science-whose-knowledge/#bookTabs=1 [Google Scholar] [Crossref]

6. Hyde, J. S., & Mezulis, A. H. (2020). Gender differences in depression: Biological, affective, cognitive, and sociocultural factors. Harvard Review of Psychiatry, 28(1), 4–13. https://doi.org/10.1097/hrp.0000000000000230 [Google Scholar] [Crossref]

7. Kautzky, A., Dold, M., Bartova, L., Spies, M., Kranz, G. S., Souery, D., Montgomery, S., Mendlewicz, J., Zohar, J., Fabbri, C., Serretti, A., Lanzenberger, R., Dikeos, D., Rujescu, D., & Kasper, S. (2019). Clinical predictors of treatment resistant depression: Replication results from the European multicenter study. European Neuropsychopharmacology, 29, S61–S62. https://doi.org/10.1016/j.euroneuro.2018.11.1038 [Google Scholar] [Crossref]

8. Marcus, S. M., Young, E. A., Kerber, K. B., Kornstein, S., Farabaugh, A. H., Mitchell, J., Wisniewski, S. R., Balasubramani, G. K., Trivedi, M. H., & Rush, A. J. (2005). Gender differences in depression: Findings from the STAR*D study. Journal of Affective Disorders, 87(2-3), 141–150. https://doi.org/10.1016/j.jad.2004.09.008 [Google Scholar] [Crossref]

9. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A Survey on Bias and Fairness in Machine Learning. ACM Computing Surveys, 54(6), 1–35. https://doi.org/10.1145/3457607 [Google Scholar] [Crossref]

10. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342 [Google Scholar] [Crossref]

11. Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., & Moher, D. (2021). Updating guidance for reporting systematic reviews: development of the PRISMA 2020 statement. Journal of Clinical Epidemiology, 134(134), 103–112. https://doi.org/10.1016/j.jclinepi.2021.02.003 [Google Scholar] [Crossref]

12. Raj, A., Ali, Z., Chaudhary, S., Bali, K. K., & Sharma, A. (2024). Depression Detection Using BERT on Social Media Platforms. 2022 IEEE International Conference on Artificial Intelligence in Engineering and Technology (IICAIET), 228–233. https://doi.org/10.1109/iicaiet62352.2024.10730329 [Google Scholar] [Crossref]

13. Schnack, H. G., & Kahn, R. S. (2016). Detecting Neuroimaging Biomarkers for Psychiatric Disorders: Sample Size Matters. Frontiers in Psychiatry, 7. https://doi.org/10.3389/fpsyt.2016.00050 [Google Scholar] [Crossref]

14. Shatte, A. B. R., Hutchinson, D. M., & Teague, S. J. (2019). Machine learning in mental health: a scoping review of methods and applications. Psychological Medicine, 49(09), 1426–1448. https://doi.org/10.1017/s0033291719000151 [Google Scholar] [Crossref]

15. Thakkar, A., Gupta, A., & De Sousa, A. (2024). Artificial intelligence in positive mental health: a narrative review. Frontiers in Digital Health, 6(1280235). https://doi.org/10.3389/fdgth.2024.1280235 [Google Scholar] [Crossref]

16. Wolff, R. F., Moons, K. G. M., Riley, R. D., Whiting, P. F., Westwood, M., Collins, G. S., Reitsma, J. B., Kleijnen, J., & Mallett, S. (2019). PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Annals of Internal Medicine, 170(1), 51. https://doi.org/10.7326/m18-1376 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles