Response Stability of Large Language Models Under Meaning-Preserving Prompt Variation

Authors

Idrees

College of Foreign Languages and Cultures, Chengdu University of Technology, Erqianqiao #1, Chengdu, Sichuan 610059 (China)

Liu Yongzhi

College of Foreign Languages and Cultures, Chengdu University of Technology, Erqianqiao #1, Chengdu, Sichuan 610059 (China)

Article Information

DOI: 10.47772/IJRISS.2026.100800423

Subject Category: Education

Volume/Issue: 10/8 | Page No: 6605-6628

Publication Timeline

Submitted: 2026-08-19

Accepted: 2026-08-24

Published: 2026-09-07

Abstract

Single-prompt assessments provide limited evidence about whether task-relevant response properties remain stable when the same communicative intention is reformulated. This study examined response stability under controlled, meaning-preserving English prompt variation using 648 responses from 12 researcher-selected argumentative policy and education topics, six prompt variants, three accessed consumer systems, and three repeated runs. The analysis combined response-length screening, exact-duplication and cross-topic reuse detection, latent semantic analysis (LSA), structural instruction checks, bootstrap confidence intervals, and matched non-parametric comparisons. Under a fixed 100-dimensional LSA configuration, aggregate semantic stability was .975 for ChatGPT (95% bootstrap CI [.965, .985]), .968 for Claude [.948, .985], and .860 for Gemini [.791, .924]. ChatGPT and Claude showed comparably high semantic stability, whereas the accessed Gemini configuration was lower and more variable. All three run-specific Friedman tests yielded χ²(2) = 24.000, p = 6.14 × 10⁻⁶, Kendall’s W = 1.000; this value indicates consistent within-run rank ordering across the 12 topics, not model superiority, and includes ties produced by exact repetition. ChatGPT and Claude met the experimental 180–250-word instruction in 100% of responses, compared with 33.3% for Gemini. Repetition complicated interpretation: Claude returned byte-identical six-variant blocks in 66.7% of run-topic blocks and Gemini in 33.3%, while Gemini Run 2 produced only 25 distinct texts across 72 responses without any complete block. LSA estimates also varied with representation settings, so model ordering based on small semantic-score differences was not robust. Only three of nine runs retained version and tier information, those documented configurations used unequal access tiers, and other session and inference conditions could not be reconstructed. The findings therefore describe the accessed consumer configurations under this task design, not controlled vendor-level or model-family differences. The study supports a multidimensional operational profile that reports semantic preservation, instruction following, topical relevance, structural adaptation, and repetition as related but distinct dimensions.

Keywords

large language models, response stability, prompt variation, semantic similarity, instruction following

Downloads

References

1. Bommasani, R., Liang, P., & Lee, T. (2023). Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525(1), 140–146. https://doi.org/10.1111/nyas.15007 [Google Scholar] [Crossref]

2. Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., Ye, W., Zhang, Y., Chang, Y., Yu, P. S., Yang, Q., & Xie, X. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), Article 39, 1–45. https://doi.org/10.1145/3641289 [Google Scholar] [Crossref]

3. Chatterjee, A., Renduchintala, H. S. V. N. S. K., Bhatia, S., & Chakraborty, T. (2024). POSIX: A prompt sensitivity index for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 14550–14565). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.852 [Google Scholar] [Crossref]

4. Davison, A. C., & Hinkley, D. V. (1997). Bootstrap methods and their application. Cambridge University Press. https://doi.org/10.1017/CBO9780511802843 [Google Scholar] [Crossref]

5. Deerwester, S., Dumais, S. T., Furnas, G. W., Landauer, T. K., & Harshman, R. (1990). Indexing by latent semantic analysis. Journal of the American Society for Information Science, 41(6), 391–407. https://doi.org/10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9 [Google Scholar] [Crossref]

6. Fu, T., & Barez, F. (2025). Same question, different words: A latent adversarial framework for prompt robustness. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 31464–31481). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-main.1595 [Google Scholar] [Crossref]

7. Fu, W., Wei, B., Hu, J., Cai, Z., & Liu, J. (2024). QGEval: Benchmarking multi-dimensional evaluation for question generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 11783–11803). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.658 [Google Scholar] [Crossref]

8. Gan, C., & Mori, T. (2023). Sensitivity and robustness of large language models to prompt template in Japanese text classification tasks. In Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation (pp. 1–11). Association for Computational Linguistics. https://aclanthology.org/2023.paclic-1.1/ [Google Scholar] [Crossref]

9. Gehrmann, S., Adewumi, T., Aggarwal, K., Ammanamanchi, P. S., Anuoluwapo, A., Bosselut, A., Chandu, K. R., Clinciu, M.-A., Das, D., Dhole, K. D., Du, W., Durmus, E., Dušek, O., Emezue, C., Gangal, V., Garbacea, C., Hashimoto, T., Hou, Y., Jernite, Y., … Zhou, J. (2021). The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the First Workshop on Natural Language Generation, Evaluation, and Metrics (pp. 96–120). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.gem-1.10 [Google Scholar] [Crossref]

10. Hämäläinen, M., & Alnajjar, K. (2021). Human evaluation of creative NLG systems: An interdisciplinary survey on recent papers. In Proceedings of the First Workshop on Natural Language Generation, Evaluation, and Metrics (pp. 84–95). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.gem-1.9 [Google Scholar] [Crossref]

11. Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89. https://doi.org/10.1080/19312450709336664 [Google Scholar] [Crossref]

12. Hua, A., Tang, K., Gu, C., Gu, J., Wong, E., & Qin, Y. (2025). Flaw or artifact? Rethinking prompt sensitivity in evaluating LLMs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 19889–19899). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-main.1006 [Google Scholar] [Crossref]

13. Kerby, D. S. (2014). The simple difference formula: An approach to teaching nonparametric correlation. Comprehensive Psychology, 3, Article 11.IT.3.1. https://doi.org/10.2466/11.IT.3.1 [Google Scholar] [Crossref]

14. Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, Article 863. https://doi.org/10.3389/fpsyg.2013.00863 [Google Scholar] [Crossref]

15. Leidinger, A., van Rooij, R., & Shutova, E. (2023). The language of prompting: What linguistic properties make a prompt successful? In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 9210–9232). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.618 [Google Scholar] [Crossref]

16. Lior, G., Habba, E., Levy, S., Caciularu, A., & Stanovsky, G. (2025). ReliableEval: A recipe for stochastic LLM evaluation via method of moments. In Findings of the Association for Computational Linguistics: EMNLP 2025. Association for Computational Linguistics. https://aclanthology.org/2025.findings-emnlp.594/ [Google Scholar] [Crossref]

17. Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 2511–2522). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.153 [Google Scholar] [Crossref]

18. Lu, Y., Bartolo, M., Moore, A., Riedel, S., & Stenetorp, P. (2022). Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 8086–8098). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.556 [Google Scholar] [Crossref]

19. Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., & Stanovsky, G. (2024). State of what art? A call for multi-prompt LLM evaluation. Transactions of the Association for Computational Linguistics, 12, 933–949. https://doi.org/10.1162/tacl_a_00681 [Google Scholar] [Crossref]

20. Moradi, M., & Samwald, M. (2021). Evaluating the robustness of neural language models to input perturbations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 1558–1570). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.117 [Google Scholar] [Crossref]

21. Nalbandyan, G., Shahbazyan, R., & Bakhturina, E. (2025). SCORE: Systematic consistency and robustness evaluation for large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Industry Track (pp. 470–484). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-industry.39 [Google Scholar] [Crossref]

22. Qiang, Y., Nandi, S., Mehrabi, N., Ver Steeg, G., Kumar, A., Rumshisky, A., & Galstyan, A. (2024). Prompt perturbation consistency learning for robust language models. In Findings of the Association for Computational Linguistics: EACL 2024 (pp. 1357–1370). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-eacl.91 [Google Scholar] [Crossref]

23. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410 [Google Scholar] [Crossref]

24. Schmidtova, P., Mahamood, S., Balloccu, S., Dušek, O., Gatt, A., Gkatzia, D., Howcroft, D. M., Platek, O., & Sivaprasad, A. (2024). Automatic metrics in natural language generation: A survey of current evaluation practices. In Proceedings of the 17th International Natural Language Generation Conference (pp. 557–583). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.inlg-main.44 [Google Scholar] [Crossref]

25. Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design, or: How I learned to start worrying about prompt formatting. In International Conference on Learning Representations. https://openreview.net/forum?id=RIu5lyNXjT [Google Scholar] [Crossref]

26. Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., … Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172–180. https://doi.org/10.1038/s41586-023-06291-2 [Google Scholar] [Crossref]

27. Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Amin, M., Hou, L., Clark, K., Pfohl, S. R., Cole-Lewis, H., Neal, D., Rashid, Q. M., Schaekermann, M., Wang, A., Dash, D., Chen, J. H., Shah, N. H., Lachgar, S., Mansfield, P. A., … Natarajan, V. (2025). Toward expert-level medical question answering with large language models. Nature Medicine, 31(3), 943–950. https://doi.org/10.1038/s41591-024-03423-7 [Google Scholar] [Crossref]

28. Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., … Wu, Z. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research. https://openreview.net/forum?id=uyTL5Bvosj [Google Scholar] [Crossref]

29. Stureborg, R., Alikaniotis, D., & Suhara, Y. (2024). Characterizing the confidence of large language model-based automatic evaluation metrics. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 76–89). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.eacl-short.9 [Google Scholar] [Crossref]

30. Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29, 1930–1940. https://doi.org/10.1038/s41591-023-02448-8 [Google Scholar] [Crossref]

31. Vu, T., Krishna, K., Alzubi, S., Tar, C., Faruqui, M., & Sung, Y.-H. (2024). Foundational autoraters: Taming large language models for better automatic evaluation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 17086–17105). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.949 [Google Scholar] [Crossref]

32. Wahle, J. P., Ruas, T., Xu, Y., & Gipp, B. (2024). Paraphrase types elicit prompt engineering capabilities. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 11004–11033). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.617 [Google Scholar] [Crossref]

33. Wang, Y., Chen, L., Cai, S., Xu, Z., & Zhao, Y. (2024). Revisiting automated evaluation for long-form table question answering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 14663–14677). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.815 [Google Scholar] [Crossref]

34. Watts, I., Gumma, V., Yadavalli, A., Seshadri, V., Swaminathan, M., & Sitaram, S. (2024). PARIKSHA: A large-scale investigation of human-LLM evaluator agreement on multilingual and multi-cultural data. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 7900–7932). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.451 [Google Scholar] [Crossref]

35. Webson, A., & Pavlick, E. (2022). Do prompt-based models really understand the meaning of their prompts? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 2300–2344). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.naacl-main.167 [Google Scholar] [Crossref]

36. Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr [Google Scholar] [Crossref]

37. Zhelezniak, V., Savkov, A., Shen, A., & Hammerla, N. (2019). Correlation coefficients and semantic textual similarity. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 951–962). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1100 [Google Scholar] [Crossref]

38. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36: Datasets and Benchmarks Track. https://papers.nips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html [Google Scholar] [Crossref]

39. Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., & Hou, L. (2023). Instruction-following evaluation for large language models. arXiv. https://arxiv.org/abs/2311.07911 [Google Scholar] [Crossref]

40. Zhuo, J., Zhang, S., Fang, X., Duan, H., Lin, D., & Chen, K. (2024). ProSA: Assessing and understanding the prompt sensitivity of LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 1950–1976). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.108 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles