Statistical Comparison of Initial and Confirmatory (Final) Diagnostic Screening Test Results for a Condition in a Population
Authors
Precious Onyedikachi Ibeakuzie
Department of Statistics, Faculty of Physical Sciences, Nnamdi Azikiwe University, Awka (Nigeria)
Department of Statistics, Faculty of Physical Sciences, Nnamdi Azikiwe University, Awka (Nigeria)
Department of Statistics, Faculty of Physical Sciences, Nnamdi Azikiwe University, Awka (Nigeria)
Consultant Haematologist, Department of Haematology and Blood Banking, Nnamdi Azikiwe University Teaching Hospital, Nnewi (Nigeria)
Consultant Ophthalmologist, Department of Ophthalmology, Nnamdi Azikiwe University Teaching Hospital, Nnewi (Nigeria)
Article Information
DOI: 10.47772/IJRISS.2026.100800249
Subject Category: Mathematics
Volume/Issue: 10/8 | Page No: 3683-3698
Publication Timeline
Submitted: 2026-08-21
Accepted: 2026-08-26
Published: 2026-09-01
Abstract
Background and Objectives: Diagnostic screening tests are routinely used to detect a condition in a population. When the same subjects undergo an initial screen and a confirmatory test, the consistency of the two sets of results must be assessed. Existing methods answer this only in part: the McNemar test detects marginal heterogeneity without quantifying the direction of discordance, while Cohen’s kappa measures overall agreement without isolating directional asymmetry. This paper proposes a nonparametric method that explicitly addresses directional discordance between initial and confirmatory screening results.
Methods: Trichotomous scores of 1, 0 and -1 are assigned to subjects according to their response patterns across the two occasions. An aggregate statistic W is derived with its variance, and a chi-square statistic with one degree of freedom is obtained. The method is compared analytically with the McNemar test and Cohen’s kappa, and a Monte Carlo study of 10,000 replications evaluates its Type I error rate and power against the McNemar test for sample sizes from 50 to 2000.
Results: The method was applied to published data on 305 hospitalised patients in whom conjunctival pallor was the initial screen for severe anaemia and laboratory haemoglobin below 7 g/dL the confirmatory test. The proportions shifting from positive-to-negative and negative-to-positive were 27.9% and 0.7%, a 42-fold imbalance. The chi-square statistic was 106.950 (df=1, p<0.001) against 79.184 for the McNemar test, while Cohen’s kappa of 0.230 disclosed neither the direction nor the magnitude of the imbalance. Simulation confirmed nominal Type I error for n≥100 and power at least matching McNemar.
Conclusions: Conjunctival pallor is a dependable rule-out but unreliable rule-in sign for severe anaemia: it rarely misses a case yet over-classifies heavily, so haemoglobin confirmation remains indispensable. The method fills a gap in paired categorical data analysis by explicitly quantifying directional discordance.
Keywords
diagnostic screening; directional discordance; trichotomous scoring; McNemar test; Cohen’s kappa
Downloads
References
1. Pepe, M. S. (2003). The statistical evaluation of medical tests for classification and prediction. Oxford University Press. [Google Scholar] [Crossref]
2. Zhou, X. H., Obuchowski, N. A., & McClish, D. K. (2011). Statistical methods in diagnostic medicine (2nd ed.). John Wiley & Sons. https://doi.org/10.1002/9780470906514 [Google Scholar] [Crossref]
3. McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2), 153-157. https://doi.org/10.1007/BF02295996 [Google Scholar] [Crossref]
4. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46. https://doi.org/10.1177/001316446002000104 [Google Scholar] [Crossref]
5. Butt, Z., Ashfaq, U., Sherazi, S. F. H., Jan, N. U., & Shahbaz, U. (2010). Diagnostic accuracy of “pallor” for detecting mild and severe anaemia in hospitalized patients. Journal of the Pakistan Medical Association, 60(9), 762-765. [Google Scholar] [Crossref]
6. Stuart, A. (1955). A test for homogeneity of the marginal distributions in a two-way classification. Biometrika, 42(3-4), 412-416. https://doi.org/10.1093/biomet/42.3-4.412 [Google Scholar] [Crossref]
7. Bowker, A. H. (1948). A test for symmetry in contingency tables. Journal of the American Statistical Association, 43(244), 572-574. https://doi.org/10.1080/01621459.1948.10483284 [Google Scholar] [Crossref]
8. Bhapkar, V. P. (1966). A note on the equivalence of two test criteria for hypotheses in categorical data. Journal of the American Statistical Association, 61(313), 228-235. https://doi.org/10.1080/01621459.1966.10502021 [Google Scholar] [Crossref]
9. Agresti, A. (2002). Categorical data analysis (2nd ed.). John Wiley & Sons. https://doi.org/10.1002/0471249688 [Google Scholar] [Crossref]
10. Fleiss, J. L., Levin, B., & Paik, M. C. (2003). Statistical methods for rates and proportions (3rd ed.). John Wiley & Sons. https://doi.org/10.1002/0471445428 [Google Scholar] [Crossref]
11. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159-174. https://doi.org/10.2307/2529310 [Google Scholar] [Crossref]
12. Senn, S. (2002). Cross-over trials in clinical research (2nd ed.). John Wiley & Sons. https://doi.org/10.1002/0470854596 [Google Scholar] [Crossref]
13. Jones, B., & Kenward, M. G. (2014). Design and analysis of cross-over trials (3rd ed.). CRC Press. [Google Scholar] [Crossref]
14. Bain, B. J., Bates, I., & Laffan, M. A. (2017). Dacie and Lewis practical haematology (12th ed.). Elsevier. [Google Scholar] [Crossref]
15. Bland, J. M., & Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327(8476), 307-310. https://doi.org/10.1016/S0140-6736(86)90837-8 [Google Scholar] [Crossref]
16. Conover, W. J. (1999). Practical nonparametric statistics (3rd ed.). John Wiley & Sons. [Google Scholar] [Crossref]
17. Hollander, M., Wolfe, D. A., & Chicken, E. (2014). Nonparametric statistical methods (3rd ed.). John Wiley & Sons. [Google Scholar] [Crossref]
18. Oyeka, I. C. A., & Ebuh, G. U. (2012). Modified Wilcoxon signed-rank test. Open Journal of Statistics, 2(2), 172-176. https://doi.org/10.4236/ojs.2012.22019 [Google Scholar] [Crossref]
19. Chalco, J. P., Huicho, L., Alamo, C., Carreazo, N. Y., & Bada, C. A. (2005). Accuracy of clinical pallor in the diagnosis of anaemia in children: A meta-analysis. BMC Pediatrics, 5, 46. https://doi.org/10.1186/1471-2431-5-46 [Google Scholar] [Crossref]
20. Kalantri, A., Karambelkar, M., Joshi, R., Kalantri, S., & Jajoo, U. (2010). Accuracy and reliability of pallor for detecting anaemia: A hospital-based diagnostic accuracy study. PLoS ONE, 5(1), e8545. https://doi.org/10.1371/journal.pone.0008545 [Google Scholar] [Crossref]
21. Ogunfowokan, O., Ogunfowokan, B. A., & Nwajei, A. I. (2020). Sensitivity and specificity of malaria rapid diagnostic test (mRDT CareStat) compared with microscopy amongst under five children attending a primary care clinic in southern Nigeria. African Journal of Primary Health Care & Family Medicine, 12(1), a2212. https://doi.org/10.4102/phcfm.v12i1.2212 [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Interplay of Students’ Emotional Intelligence and Attitude toward Mathematics on Performance in Grade 10 Algebra
- Numerical Simulation of Fitzhugh-Nagumo Dynamics Using a Finite Difference-Based Method of Lines
- Fixed Point Theorem in Controlled Metric Spaces
- Usage of Moving Average to Heart Rate, Blood Pressure and Blood Sugar
- Exploring Algebraic Topology and Homotopy Theory: Methods, Empirical Data, and Numerical Examples