Say My Name: How Declaring Black Identity Triggers the Safety Filters that Writing Black Does Not
Authors
CIS/IT Department, Purdue University Global (USA)
Article Information
DOI: 10.51244/IJRSI.2026.1313CS015
Subject Category: Language
Volume/Issue: 13/13 | Page No: 186-194
Publication Timeline
Submitted: 2026-06-18
Accepted: 2026-06-24
Published: 2026-07-01
Abstract
This paper presents an independent secondary validation of Haq & Saldías (2026), which reported that large language models refuse requests from explicitly Black identified users at rates 7.5 to 8.27 percentage points above those observed for White identified users, and that African American Vernacular English (AAVE) dialect use nearly eliminates this penalty a pattern the authors term the “dialect jailbreak.” All validation is conducted without new model inference, deriving exclusively from the study's published statistics and the publicly available BOLD dataset (Dhamala et al., 2021). Three findings are reported. First, the paper's core refusal rate claims and the 97.82% soft refusal figure are arithmetically confirmed, and the primary Average Marginal Effects are independently replicated via two proportion z tests (Black vs. White Explicit: +8.27pp, p < .001). Second, the conclusion level claim of a "400% increase in odds" requires substantive qualification: the figure reflects a GLMM conditional odds ratio inflated by extraordinary prompt level random effect variance (σ² = 100.44, ICC = .97), diverging approximately 5-fold from the raw marginal odds ratio of 1.56 a distinction the original paper does not address. Third, a novel additive decomposition analysis reveals that the AAVE dialect jailbreak is sub additive, with the observed refusal rate falling 7.41 percentage points below what independent race and dialect penalties would additively predict, indicating synergistic rather than sequential safety filter degradation. Taken together, these findings suggest that keyword-gated alignment mechanisms are doubly brittle: susceptible to explicit identity bypass through lexical substitution, and to multi-signal compounding that exceeds what single signal analyses anticipate. An interactive validation dashboard and all derived data files are publicly available at https://github.com/DrMarioDeSean411-arch/say-my-name-llm-bias
Keywords
large language model alignment, racial bias
Downloads
References
1. Booker, M. D. (2022). Empowering, engaging, and equipping technology through acceptance: A quantitative study utilizing the modified technology acceptance model (mTAM) to explore facial-recognition software acceptance in the smart policing initiative (SPI) [Doctoral dissertation, University of the Cumberlands]. ResearchGate. https://doi.org/10.13140/RG.2.2.13108.77441 [Google Scholar] [Crossref]
2. Booker, M. D. (2026). Say my name: Interactive validation dashboard — LLM refusal bias study [Software]. GitHub. https://github.com/DrMarioDeSean411-arch/say-my-name-llm-bias [Google Scholar] [Crossref]
3. Crenshaw, K. (1989). Demarginalizing the intersection of race and sex: A Black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. University of Chicago Legal Forum, 1989(1), 139–167. [Google Scholar] [Crossref]
4. Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K. W., & Gupta, R. (2021). BOLD: Dataset and metrics for measuring biases in open-ended language generation. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 862–872. https://doi.org/10.1145/3442188.3445924 [Google Scholar] [Crossref]
5. Haq, I., & Saldías, B. (2026). Dialect vs. demographics: Quantifying LLM bias from implicit linguistic signals vs. explicit user profiles. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26). https://doi.org/10.1145/3805689.3812419 [Google Scholar] [Crossref]
6. Neuhaus, J. M., Kalbfleisch, J. D., & Hauck, W. W. (1991). A comparison of cluster-specific and population-averaged approaches for analyzing correlated binary data. International Statistical Review, 59(1), 25–35. [Google Scholar] [Crossref]
7. Sheng, E., Chang, K.-W., Natarajan, P., & Peng, N. (2019). The woman worked as a babysitter: On biases in language generation. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 3407–3412. https://doi.org/10.18653/v1/D19-1339 [Google Scholar] [Crossref]
8. Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. Proceedings of the 8th International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Evaluating the Impacts of Mind Mapping Strategy on Developing EFL Students’ Critical Reading Skills
- Significance of Reading Instructions for Language Improvement in Children with Down Syndrome
- Prenasalised Consonants in Liangmai
- Metadiscourse Matters: Definitions, Models, and Advantages for ESL/ EFL Writing
- Blank Minds and Stuck Voices: Understanding and Addressing Cognitive Anxiety in High-Stakes ESL Speaking Tests