Similarity-Based Target Leakage in Case-Based Reasoning for Stress Risk Screening: A Cautionary Demonstration

Authors

Jonas Niyitegeka

Kigali Independent University, Kigali (Rwanda)

Anaclet Ukurikiyeyezu

African Institute for Mathematical Sciences (AIMS), Kigali (Rwanda)

Shaloom Niyibizi

Kigali Independent University, Kigali (Rwanda)

Article Information

DOI: 10.51244/IJRSI.2026.1309000049

Subject Category: Science

Volume/Issue: 13/9 | Page No: 625-639

Publication Timeline

Submitted: 2026-09-09

Accepted: 2026-09-14

Published: 2026-10-03

Abstract

Early identification of elevated stress may support preventive mental health interventions, particularly in low-resource settings. This study evaluates a Case-Based Reasoning (CBR) approach alongside conventional Machine Learning (ML) classifiers for stress risk screening using an anonymized COVID-19 quarantine survey dataset (N=824) containing twelve categorical predictors and three approximately balanced Growing_Stress classes. Six supervised learning algorithms (Random Forest, k-Nearest Neighbours, XGBoost, Decision Tree, Support Vector Machine, and Logistic Regression) were evaluated using stratified cross-validation, while a CBR model based on Jaccard similarity and mode-based revision was assessed using a stratified hold-out protocol. Under a leakage-free experimental design restricting similarity computation to predictor attributes only, both CBR and ML models achieved performance close to random classification (accuracy 0.33–0.38; macro-AUC ≈0.50), indicating limited predictive information in the available behavioural features. A controlled ablation experiment intentionally including the target attribute in the CBR similarity computation increased accuracy to 0.94–0.99 and AUC to 1.00, without changing the dataset, model structure, or evaluation procedure. This demonstrates that similarity-based retrieval can be vulnerable to target leakage when solution attributes are unintentionally incorporated during retrieval, producing misleadingly high performance estimates. The contribution of this study is twofold: it provides an empirical finding that the available behavioural features offered insufficient predictive information for stress classification in the study dataset, while also presenting a reproducible demonstration of how target leakage can be identified and prevented in similarity-based retrieval systems. These findings highlight the importance of explicit predictor–target separation and rigorous validation protocols for developing trustworthy AI-based decision-support systems in mental health applications. An additional identical-fold cross-validation, a majority-class baseline, and label-permutation null distributions further confirm that neither CBR nor the machine learning baselines exceed chance-level performance, and that the leakage-driven accuracy gain persists unchanged even when the leaked label is replaced with an arbitrary, permuted one, confirming the effect is a structural retrieval artefact rather than a genuine relationship in the data.

Keywords

Case-Based Reasoning, Machine Learning, Growing Stress, Target Leakage, COVID-19 Quarantine Survey

Downloads

References

1. G. B. D. Mental and D. Collaborators, “Global , regional , and national burden of 12 mental disorders in 204 countries and territories , 1990 – 2019 : a systematic analysis for the Global Burden of Disease Study 2019,” The Lancet Psychiatry, vol. 9, no. 2, pp. 137–150, 2022, doi: 10.1016/S2215-0366(21)00395-3. [Google Scholar] [Crossref]

2. G. M. Slavich and M. R. Irwin, “From Stress to Inflammation and Major Depressive Disorder : A Social Signal Transduction Theory of Depression,” vol. 140, no. 3, 2014, doi: 10.1037/a0035302. [Google Scholar] [Crossref]

3. S. Graham, C. Depp, E. E. Lee, C. Nebeker, X. Tu, H.-C. Kim, and D. V. Jeste, “Artificial Intelligence for Mental Health and Mental Illnesses: An Overview,” Current Psychiatry Reports, vol. 21, no. 11, p. 116, 2019, doi: 10.1007/s11920-019-1094-0. [Google Scholar] [Crossref]

4. S. Kaufman, S. Rosset, C. Perlich, and O. R. I. Stitelman, “Leakage in Data Mining : Formulation , Detection , and Avoidance,” vol. 6, no. 4, pp. 1–21, 2012, doi: 10.1145/2382577.2382579. [Google Scholar] [Crossref]

5. A. Aamodt, “Case-Based Reasoning : Foundational Issues , Methodological Variations , and System Approaches,” vol. 7, no. 1, pp. 39–59, 1994. [Google Scholar] [Crossref]

6. H. Burkhard, “2 . Extending some Concepts of CBR – Foundations of Case Retrieval Nets,” no. Chapter 3. [Google Scholar] [Crossref]

7. M. M. Aldarwish, “Posts,” pp. 282–285, 2017, doi: 10.1109/ISADS.2017.41. [Google Scholar] [Crossref]

8. A. B. R. Shatte, D. M. Hutchinson, and S. J. Teague, “Machine learning in mental health : a scoping review of methods and applications,” 2019. [Google Scholar] [Crossref]

9. “Machine Learning Algorithms for Depression : Diagnosis ,” pp. 1–20, 2022. [Google Scholar] [Crossref]

10. S. Wshah, C. Skalka, and M. Price, “Predicting Posttraumatic Stress Disorder Risk : A Machine Learning Approach Corresponding Author :,” vol. 6, 2019, doi: 10.2196/13946. [Google Scholar] [Crossref]

11. A. M. Y. Tai et al., “Arti fi cial Intelligence In Medicine Machine learning and big data : Implications for disease modeling and therapeutic discovery in psychiatry,” Artif. Intell. Med., vol. 99, no. August, p. 101704, 2019, doi: 10.1016/j.artmed.2019.101704. [Google Scholar] [Crossref]

12. S. Kapoor and A. Narayanan, “Article Leakage and the reproducibility crisis in machine- learning-based science Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, p. 100804, 2023, doi: 10.1016/j.patter.2023.100804. [Google Scholar] [Crossref]

13. J. L. Kolodner, “Educational Implications of Analogy”. [Google Scholar] [Crossref]

14. Y. Kang, S. Krishnaswamy, and A. Zaslavsky, “A Retrieval Strategy for Case-Based Reasoning Using Similarity and Association Knowledge,” vol. 44, no. 4, pp. 473–487, 2014. [Google Scholar] [Crossref]

15. G. Albora, M. Straccamore, and A. Zaccaria, “Machine learning-inspired similarity measure to forecast M & A from patent data,” pp. 1–19, 2026, doi: 10.1371/journal.pone.0341010. [Google Scholar] [Crossref]

16. A. C. Butler, J. E. Chapman, E. M. Forman, and A. T. Beck, “The empirical status of cognitive-behavioral therapy : A review of meta-analyses,” vol. 26, pp. 17–31, 2006, doi: 10.1016/j.cpr.2005.07.003. [Google Scholar] [Crossref]

17. N. Amin, I. Salehin, M. A. Baten, and R. Al Noman, “RHMCD-20 dataset: Identify rapid human mental health depression during quarantine life using machine learning,” Data in Brief, vol. 54, p. 110376, 2024, doi: 10.1016/j.dib.2024.110376. [Google Scholar] [Crossref]

18. R. R. Sokal and C. D. Michener, “A statistical method for evaluating systematic relationships,” University of Kansas Science Bulletin, vol. 38, pp. 1409–1438, 1958. [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles