Harnessing AI for CEFR-Aligned Writing Assessment: A Study of a Customised GPT

Authors

Chua Wei Chuan

SBP Integrasi Sabak Bernam, Selangor (Malaysia)

Melor Md Yunus

Faculty of Education, Universiti Kebangsaan Malaysia (Malaysia)

Harwati Hashim

Faculty of Education, Universiti Kebangsaan Malaysia (Malaysia)

Article Information

DOI: 10.47772/IJRISS.2026.100400326

Subject Category: Education

Volume/Issue: 10/4 | Page No: 4485-4501

Publication Timeline

Submitted: 2026-04-08

Accepted: 2026-04-14

Published: 2026-05-08

Abstract

Artificial intelligence (AI) tools, specifically ChatGPT, has garnered considerable interest in the field of English as a Second Language (ESL) as demand grows in recent times. The introduction of ChatGPT has transformed the assessment landscape, thus marking the beginning of a new era in AI-assisted assessment. While existing research has largely addressed the comparison of ChatGPT and human raters with regards to assessment, limited study has focused on customised GPTs and the context of Common European Framework of References for Languages (CEFR). To address this gap, this paper investigates the effectiveness of a customised GPT as a formative and summative assessment tool in a CEFR-aligned written task pitched at B2 CEFR level. Adopting a quantitative research design, the respondents’ attitudes towards using the GPT for formative assessment were examined through a questionnaire administered to 31 English teachers. In parallel, the scores assigned by the GPT and the teachers for the same writing tasks were compared via inter-rater reliability analysis. Findings revealed that the respondents hold a generally positive view regarding the effectiveness of GPT in providing formative feedback. However, the results also indicated that the GPT demonstrates a moderate level of agreement with the teacher scores in most assessment constructs. The data further emphasized the need of prompt engineering in developing the GPT to be an effective formative and summative assessment assistant for teachers. This paper concludes by discussing the practical implications of employing the customised GPT in assessment, thereby contributing to the discourse on AI-assisted assessment.

Keywords

Assessment, Writing, AI, Customised GPT, CEFR

Downloads

References

1. Abduljawad, S. A. (2024). Investigating the impact of ChatGPT as an AI tool on ESL writing: Prospects and challenges in Saudi Arabian higher education. International Journal of Computer-Assisted Language Learning and Teaching, 14(1), 1-19. https://doi.org/10.4018/IJCALLT.367276 [Google Scholar] [Crossref]

2. Al-khresheh, M. H. (2024). Bridging technology and pedagogy from a global lens: Teachers’ perspectives on integrating ChatGPT in English language teaching. Computers and Education: Artificial Intelligence, 6, Article 100218. https://doi.org/10.1016/j.caeai.2024.100218 [Google Scholar] [Crossref]

3. Allehyani, S. H., & Algamdi, M. A. (2023). Digital competences: Early childhood teachers' beliefs and perceptions of ChatGPT application in teaching English as a Second Language (ESL). International Journal of Learning, Teaching and Educational Research, 22(11), 343-363. https://doi.org/10.26803/ijlter.22.11.18 [Google Scholar] [Crossref]

4. Alm, A., & Ohashi, L. (2024). A worldwide study on language educators’ initial response to ChatGPT. Technology in Language Teaching & Learning, 6(1), 1141. https://doi.org/10.29140/tltl.v6n1.1141 [Google Scholar] [Crossref]

5. Almashy, A., Ahmed, A., Jamshed, M., Ansari, M., Banu, S., & Warda, W. (2024). Analyzing the impact of CALL tools on English learners’ writing skills: A comparative study of errors correction. World Journal of English Language, 14(6), p657. doi:http://dx.doi.org/10.5430/wjel.v14n6p657 [Google Scholar] [Crossref]

6. Alsaweed, W., & Aljebreen, S. (2024). Investigating the accuracy of ChatGPT as a writing error correction tool. International Journal of Computer-Assisted Language Learning and Teaching, 14(1), 1-18. https://doi.org/10.4018/IJCALLT.364847 [Google Scholar] [Crossref]

7. Anderson, J. A., & Ayaawan, A. E. (2023). Formative feedback in a writing programme at the University of Ghana. In A. Esimaje, B. van Rooy, D. Jolayemi, D. Nkemleke, & E. Klu (Eds.), African perspectives on the teaching and learning of English in higher education (pp. 197–213). Routledge. [Google Scholar] [Crossref]

8. Atasoy, A., Moslemi Nezhad Arani, S. (2025). ChatGPT: A reliable assistant for the evaluation of students’ written texts?. Educ Inf Technol 30, 20385–20415. https://doi.org/10.1007/s10639-025-13553-1 [Google Scholar] [Crossref]

9. Barkaoui, K., & Woodworth, J. (2023). An exploratory study of the construct measured by automated writing scores across task types and test occasions. Studies in Language Assessment, 12(1), 1–38. https://doi.org/10.58379/QCFS2805 [Google Scholar] [Crossref]

10. Barrett, A., & Pack, A. (2023). Not quite eye to AI: Student and teacher perspectives on the use of generative artificial intelligence in the writing process. International Journal of Educational Technology in Higher Education, 20(1), Article 59. https://doi.org/10.1186/s41239-023-00427-0 [Google Scholar] [Crossref]

11. Barrot, J. S. (2023). Using ChatGPT for second language writing: Pitfalls and potentials. Assessing Writing, 57, 100745. https://doi.org/10.1016/j.asw.2023.100745 [Google Scholar] [Crossref]

12. Beaumont, C., O’Doherty, M. & Shannon, L. (2011). Reconceptualising assessment feedback: A key to improving student learning? Studies in Higher Education, 36(6), 671–687. https://doi.org/10.1080/03075071003731135 [Google Scholar] [Crossref]

13. Bhutoria, A. (2022). Personalized education and Artificial Intelligence in the United States, China, and India: A systematic review using a Human-In-The-Loop model. Computers and Education: Artificial Intelligence, 3, 100068. https://doi.org/https://doi.org/10.1016/j.caeai.2022.100068 [Google Scholar] [Crossref]

14. Bonner, E., Lege, R., & Frazier, E. (2023). Large language model-based artificial intelligence in the language classroom: Practical ideas for teaching. Teaching English with Technology, 23(1), 23–41. https://doi.org/10.56297/BKAM1691/WIEO1749 [Google Scholar] [Crossref]

15. Branch, R. M. (2009). Instructional design: The ADDIE approach. Springer. [Google Scholar] [Crossref]

16. Cambridge Assessment English. (2020). SPM – Writing. Video. Putrajaya, Ministry of Education Malaysia. [Google Scholar] [Crossref]

17. Chua, Y. P. (2013). Mastering research statistics. McGraw-Hill Education. [Google Scholar] [Crossref]

18. Chung, J. Y., & Jeong, S.-H. (2024). Exploring the perceptions of Chinese pre-service teachers on the integration of generative AI in English language teaching: Benefits, challenges, and educational implications. Online Journal of Communication and Media Technologies, 14(4), e202457. https://doi.org/10.30935/ojcmt/15266 [Google Scholar] [Crossref]

19. Cole, W. R., Arrieux, J. P., Schwab, K., Ivins, B. J., Qashu, F. M., & Lewis, S. C. (2013). Test–retest reliability of four computerized neurocognitive assessment tools in an active duty military population. Archives of Clinical Neuropsychology, 28, 732–742. https://doi.org/10.1093/arclin/act040 [Google Scholar] [Crossref]

20. Council of Europe. (2001). Common European Framework of Reference for Languages: Language, teaching, assessment. Cambridge University Press. [Google Scholar] [Crossref]

21. Council of Europe. (2026). Historical overview of the development of the CEFR. https://www.coe.int/en/web/common-european-framework-reference-languages/history [Google Scholar] [Crossref]

22. Creswell, J. W. (2012). Educational research: Planning, conducting and evaluating quantitative and qualitative research (4th ed.). Pearson. [Google Scholar] [Crossref]

23. Crossley, A. S., Varner, L. K., Roscoe, R. D., & McNamara, D. S. (2013). Using automated indices of cohesion to evaluate an intelligent tutoring system and an automated writing evaluation system. In H. C. Lane, K. Yacef, J. Mostow, & P. Pavlik (Eds.), Artificial Intelligence in Education. AIED 2013. Lecture Notes in Computer Science (Vol. 7926, pp. 134–143). Springer. https://doi.org/10.1007/978-3-642-39112-5_28 [Google Scholar] [Crossref]

24. Curriculum Development Division. (2015). English Language Curriculum Framework Secondary. Ministry of Education Malaysia. [Google Scholar] [Crossref]

25. Dahri, N. A., Yahaya, N., Al-Rahmi, W. M., Aldraiweesh, A., Alturki, U., Almutairy, S., Shutaleva, A., & Soomro, R. B. (2024). Extended TAM based acceptance of AI-Powered ChatGPT for supporting metacognitive self-regulated learning in education: A mixed-methods study. Heliyon, 10(8), e29317. https://doi.org/10.1016/j.heliyon.2024.e29317 [Google Scholar] [Crossref]

26. Elkatmış, M. (2024). Chat GPT and Creative Writing: Experiences of Master’s Students in Enhancing. International Journal of Contemporary Educational Research, 11(3), 321-336. https://doi.org/10.52380/ijcer.2024.11.3.597 [Google Scholar] [Crossref]

27. English Language Standards and Quality Council (ELSQC). (2015). English language education reform in Malaysia: The roadmap 2015-2025. Ministry of Education Malaysia. [Google Scholar] [Crossref]

28. Examinations Syndicate. (2020). Format Pentaksiran Bahasa Inggeris Yang Dijajarkan Kepada CEFR. Examinations Syndicate. [Google Scholar] [Crossref]

29. Examinations Syndicate. (2021). Sijil Pelajaran Malaysia English Language: Instructions For Writing Examiners (To Be Used With Revised Examination). Examinations Syndicate. [Google Scholar] [Crossref]

30. Fitria, T. N. (2023). Artificial intelligence (AI) technology in OpenAI ChatGPT application: A review of ChatGPT in writing English essay. ELT Forum: Journal of English Language Teaching, 12(1), 44-58. https://doi.org/10.15294/elt.v12i1.64069 [Google Scholar] [Crossref]

31. Fu, Q.-K., Zou, D., Xie, H., & Cheng, G. (2022). A review of AWE feedback: Types, learning outcomes, and implications. Computer Assisted Language Learning, 37(1–2), 179–221. https://doi.org/10.1080/ 09588221.2022.2033787 [Google Scholar] [Crossref]

32. García-Varela, F., Nussbaum, M., Mendoza, M., Martínez-Troncoso, C., & Bekerman, Z. (2025). ChatGPT as a Stable and Fair Tool for Automated Essay Scoring. Education Sciences, 15(8), 946. https://doi.org/10.3390/educsci15080946 [Google Scholar] [Crossref]

33. Gardner, J., O’Leary, M., & Yuan, L. (2020). Artifcial intelligence in educational assessment: ‘Breakthrough? Or buncombe and ballyhoo?’ Journal of Computer Assisted Learning, 37, 1207–1216. https://doi.org/10.1111/jcal.12577 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles