Hybrid AI-Augmented Assessment in University English L2 Classrooms: A Tripartite Framework for Teacher Judgement, AI Evidence, and Learner Agency

Authors

Saif Al Baimani

Preparatory Studies Centre (PSC), University of Technology and Applied Sciences (Oman)

Article Information

DOI: 10.47772/IJRISS.2026.1026EDU0507

Subject Category: English Language

Volume/Issue: 10/26 | Page No: 6898-6909

Publication Timeline

Submitted: 2026-08-02

Accepted: 2026-08-07

Published: 2026-08-20

Abstract

Generative artificial intelligence (GAI) has entered English L2 classrooms faster than assessment frameworks have adapted, leaving teachers to make consequential evaluative decisions without theoretical guidance. Existing accounts document what automated evaluation does well; few specify how its use can remain compatible with the validity-argument tradition on which defensible assessment rests. This conceptual paper addresses that gap. Synthesising recent empirical evidence on AI-mediated writing assessment, the paper proposes the Hybrid AI-Augmented Assessment (HAIAA) model, which distributes evaluative work across three interdependent components. Automated Diagnostic Evaluation confines AI to rapid formative diagnosis of surface features; Human Interpretive Assessment reserves judgements of argument, rhetoric, and cultural appropriateness for instructors; Learner Reflective Verification makes learners' documented engagement with both feedback streams an assessable warrant of authorship. Operating as a recursive cycle rather than a sequence, the components protect the inference from score to construct while restoring the learner to the centre of the feedback process. Implications are drawn for classroom practitioners and curriculum designers, and the conditions under which the model would require empirical validation are specified.

Keywords

Generative artificial intelligence; L2 writing assessment; automated writing evaluation; higher-order thinking skills

Downloads

References

1. Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom's taxonomy of educational objectives. Longman. [Google Scholar] [Crossref]

2. Bachman, L. F., & Palmer, A. S. (2010). Language assessment in practice: Developing language assessments and justifying their use in the real world. Oxford University Press. [Google Scholar] [Crossref]

3. Barrot, J. S. (2023). Using ChatGPT for second language writing: Pitfalls and potentials. Assessing Writing, 57, 100745. https://doi.org/10.1016/j.asw.2023.100745 [Google Scholar] [Crossref]

4. Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. https://doi.org/10.1080/0969595980050102 [Google Scholar] [Crossref]

5. Carless, D., & Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43(8), 1315–1325. https://doi.org/10.1080/02602938.2018.1463354 [Google Scholar] [Crossref]

6. Casal, J. E., & Kessler, M. (2023). Can linguists distinguish between ChatGPT/AI and human writing? A study of research ethics and academic publishing. Research Methods in Applied Linguistics, 2(3), 100068. https://doi.org/10.1016/j.rmal.2023.100068 [Google Scholar] [Crossref]

7. Chapelle, C. A., Cotos, E., & Lee, J. (2015). Validity arguments for diagnostic assessment using automated writing evaluation. Language Testing, 32(3), 385–405. https://doi.org/10.1177/0265532214565386 [Google Scholar] [Crossref]

8. Chapelle, C. A., & Voss, E. (2016). 20 years of technology and language assessment in Language Learning & Technology. Language Learning & Technology, 20(2), 116–128. https://hdl.handle.net/10125/44464 [Google Scholar] [Crossref]

9. Corbin, T., Dawson, P., & Liu, D. (2025). Talk is cheap: Why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. https://doi.org/10.1080/02602938.2025.2503964 [Google Scholar] [Crossref]

10. Crompton, H., & Burke, D. (2024). The educational affordances and challenges of ChatGPT: State of the field. TechTrends, 68, 380–392. https://doi.org/10.1007/s11528-024-00939-0 [Google Scholar] [Crossref]

11. Darvin, R. (2025). The need for critical digital literacies in generative AI-mediated L2 writing. Journal of Second Language Writing, 67, 101186. https://doi.org/10.1016/j.jslw.2025.101186 [Google Scholar] [Crossref]

12. Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. Assessment & Evaluation in Higher Education, 49(7), 1005–1016. https://doi.org/10.1080/02602938.2024.2386662 [Google Scholar] [Crossref]

13. Dikli, S., & Bleyle, S. (2014). Automated essay scoring feedback for second language writers: How does it compare to instructor feedback? Assessing Writing, 22, 1–17. https://doi.org/10.1016/j.asw.2014.03.006 [Google Scholar] [Crossref]

14. Dimova, S., Yan, X., & Ginther, A. (2022). Local tests, local contexts. Language Testing, 39(3), 341–354. https://doi.org/10.1177/02655322221075014 [Google Scholar] [Crossref]

15. Ding, L., & Zou, D. (2024). Automated writing evaluation systems: A systematic review of Grammarly, Pigai, and Criterion with a perspective on future directions in the age of generative artificial intelligence. Education and Information Technologies, 29(11), 14151–14203. https://doi.org/10.1007/s10639-023-12402-3 [Google Scholar] [Crossref]

16. Escalante, J., Pack, A., & Barrett, A. (2023). AI-generated feedback on writing: Insights into efficacy and ENL student preference. International Journal of Educational Technology in Higher Education, 20(1), 57. https://doi.org/10.1186/s41239-023-00425-2 [Google Scholar] [Crossref]

17. Feng, H., Li, K., & Zhang, L. J. (2025). What does AI bring to second language writing? A systematic review (2014–2024). Language Learning & Technology, 29(1), 1–27. https://hdl.handle.net/10125/73629 [Google Scholar] [Crossref]

18. Godwin-Jones, R. (2022). Partnering with AI: Intelligent writing assistance and instructed language learning. Language Learning & Technology, 26(2), 5–24. https://hdl.handle.net/10125/73474 [Google Scholar] [Crossref]

19. Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. https://doi.org/10.3102/003465430298487 [Google Scholar] [Crossref]

20. Huang, X., Zou, D., Cheng, G., Chen, X., & Xie, H. (2023). Trends, research issues and applications of artificial intelligence in language education. Educational Technology & Society, 26(1), 112–131. https://doi.org/10.30191/ETS.202301_26(1).0009 [Google Scholar] [Crossref]

21. Hyland, K. (2026). Tomorrowland? Technical innovation and feedback on writing. RELC Journal, 57(1), 233–241. https://doi.org/10.1177/00336882251357675 [Google Scholar] [Crossref]

22. Jeon, J., Lee, S., & Choe, H. (2023). Beyond ChatGPT: A conceptual framework and systematic review of speech-recognition chatbots for language learning. Computers & Education, 206, 104898. https://doi.org/10.1016/j.compedu.2023.104898 [Google Scholar] [Crossref]

23. Jiang, L., Yu, S., & Wang, C. (2020). Second language writing instructors' feedback practice in response to automated writing evaluation: A sociocultural perspective. System, 93, 102302. https://doi.org/10.1016/j.system.2020.102302 [Google Scholar] [Crossref]

24. Jiang, Y., Hao, J., Fauss, M., & Li, C. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers? Computers & Education, 217, 105070. https://doi.org/10.1016/j.compedu.2024.105070 [Google Scholar] [Crossref]

25. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000 [Google Scholar] [Crossref]

26. Khalifa, M., & Albadawy, M. (2024). Using artificial intelligence in academic writing and research: An essential productivity tool. Computer Methods and Programs in Biomedicine Update, 5, 100145. https://doi.org/10.1016/j.cmpbup.2024.100145 [Google Scholar] [Crossref]

27. Kohnke, L., Moorhouse, B. L., & Zou, D. (2023). ChatGPT for language teaching and learning. RELC Journal, 54(2), 537–550. https://doi.org/10.1177/00336882231162868 [Google Scholar] [Crossref]

28. Kohnke, L., Zou, D., & Moorhouse, B. L. (2024). Technostress and English language teaching in the age of generative AI. Educational Technology & Society, 27(2), 306–320. https://doi.org/10.30191/ETS.202404_27(2).TP02 [Google Scholar] [Crossref]

29. Koltovskaia, S. (2020). Student engagement with automated written corrective feedback (AWCF) provided by Grammarly: A multiple case study. Assessing Writing, 44, 100450. https://doi.org/10.1016/j.asw.2020.100450 [Google Scholar] [Crossref]

30. Kuteeva, M., & Andersson, M. (2024). Diversity and standards in writing for publication in the age of AI—Between a rock and a hard place. Applied Linguistics, 45(3), 561–567. https://doi.org/10.1093/applin/amae025 [Google Scholar] [Crossref]

31. Lee, I. (2023). Problematising written corrective feedback: A Global Englishes perspective. Applied Linguistics, 44(4), 791–796. https://doi.org/10.1093/applin/amad038 [Google Scholar] [Crossref]

32. Leijten, M., & Van Waes, L. (2013). Keystroke logging in writing research: Using Inputlog to analyze and visualize writing processes. Written Communication, 30(3), 358–392. https://doi.org/10.1177/0741088313491692 [Google Scholar] [Crossref]

33. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779 [Google Scholar] [Crossref]

34. Lu, Y., Liles, X., & Ma, X. (2025). GenAI and human assessments of L2 Chinese writing: Interrater reliability and rater bias. Assessing Writing, 66, 100989. https://doi.org/10.1016/j.asw.2025.100989 [Google Scholar] [Crossref]

35. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. [Google Scholar] [Crossref]

36. Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2), 100050. https://doi.org/10.1016/j.rmal.2023.100050 [Google Scholar] [Crossref]

37. Mizumoto, A., Shintani, N., Sasaki, M., & Teng, M. F. (2024). Testing the viability of ChatGPT as a companion in L2 writing accuracy assessment. Research Methods in Applied Linguistics, 3(2), 100116. https://doi.org/10.1016/j.rmal.2024.100116 [Google Scholar] [Crossref]

38. Moorhouse, B. L., & Kohnke, L. (2024). The effects of generative AI on initial language teacher education: The perceptions of teacher educators. System, 122, 103290. https://doi.org/10.1016/j.system.2024.103290 [Google Scholar] [Crossref]

39. Moorhouse, B. L., & Wong, K. M. (2025). Generative artificial intelligence and language teaching. Cambridge University Press. https://doi.org/10.1017/9781009618823 [Google Scholar] [Crossref]

40. Ou, A. W., Khuder, B., Franzetti, S., & Negretti, R. (2024). Conceptualising and cultivating critical GAI literacy in doctoral academic writing. Journal of Second Language Writing, 66, 101156. https://doi.org/10.1016/j.jslw.2024.101156 [Google Scholar] [Crossref]

41. Pack, A., & Maloney, J. (2024). Using artificial intelligence in TESOL: Some ethical and pedagogical considerations. TESOL Quarterly, 58(2), 1007–1018. https://doi.org/10.1002/tesq.3255 [Google Scholar] [Crossref]

42. Pack, A., Maloney, J., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234 [Google Scholar] [Crossref]

43. Perkins, M., Furze, L., Roe, J., & MacVaugh, J. (2024). The Artificial Intelligence Assessment Scale (AIAS): A framework for ethical integration of generative AI in educational assessment. Journal of University Teaching and Learning Practice, 21(6). https://doi.org/10.53761/q3azde36 [Google Scholar] [Crossref]

44. Ranalli, J. (2021). L2 student engagement with automated feedback on writing: Potential for learning and issues of trust. Journal of Second Language Writing, 52, 100816. https://doi.org/10.1016/j.jslw.2021.100816 [Google Scholar] [Crossref]

45. Ranalli, J., Link, S., & Chukharev-Hudilainen, E. (2017). Automated writing evaluation for formative assessment of second language writing: Investigating the accuracy and usefulness of feedback as part of argument-based validation. Educational Psychology, 37(1), 8–25. https://doi.org/10.1080/01443410.2015.1136407 [Google Scholar] [Crossref]

46. Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119–144. https://doi.org/10.1007/BF00117714 [Google Scholar] [Crossref]

47. Sari, E. (2024). The impact of automated writing evaluation on English as a foreign language learners' writing self-efficacy, self-regulation, anxiety, and performance. Journal of Computer Assisted Learning, 40(3), 1012–1025. https://doi.org/10.1111/jcal.13004 [Google Scholar] [Crossref]

48. Saricaoglu, A., & Bilki, Z. (2025). The capacity of ChatGPT-4 for L2 writing assessment: A closer look at accuracy, specificity, and relevance. Annual Review of Applied Linguistics, 45, 253–273. https://doi.org/10.1017/S0267190525100160 [Google Scholar] [Crossref]

49. Shen, C., Shi, P., Guo, J., Xu, S., & Tian, J. (2023). From process to product: Writing engagement and performance of EFL learners under computer-generated feedback instruction. Frontiers in Psychology, 14, 1258286. https://doi.org/10.3389/fpsyg.2023.1258286 [Google Scholar] [Crossref]

50. Shi, H., & Aryadoust, V. (2024). A systematic review of AI-based automated written feedback research. ReCALL, 36(2), 187–209. https://doi.org/10.1017/S0958344023000265 [Google Scholar] [Crossref]

51. Song, C., & Song, Y. (2023). Enhancing academic writing skills and motivation: Assessing the efficacy of ChatGPT in AI-assisted language learning for EFL students. Frontiers in Psychology, 14, 1260843. https://doi.org/10.3389/fpsyg.2023.1260843 [Google Scholar] [Crossref]

52. Sotiriadou, P., Logan, D., Daly, A., & Guest, R. (2020). The role of authentic assessment to preserve academic integrity and promote skill development and employability. Studies in Higher Education, 45(11), 2132–2148. https://doi.org/10.1080/03075079.2019.1582015 [Google Scholar] [Crossref]

53. Su, Y., Lin, Y., & Lai, C. (2023). Collaborating with ChatGPT in argumentative writing classrooms. Assessing Writing, 57, 100752. https://doi.org/10.1016/j.asw.2023.100752 [Google Scholar] [Crossref]

54. Teng, M. F. (2024). A systematic review of ChatGPT for English as a foreign language writing: Opportunities, challenges, and recommendations. International Journal of TESOL Studies, 6(3), 36–57. https://doi.org/10.58304/ijts.20240304 [Google Scholar] [Crossref]

55. Voss, E., Cushing, S. T., Ockey, G. J., & Yan, X. (2023). The use of assistive technologies including generative AI by test takers in language assessment: A debate of theory and practice. Language Assessment Quarterly, 20(4–5), 520–532. https://doi.org/10.1080/15434303.2023.2258384 [Google Scholar] [Crossref]

56. Wang, Y., Huang, J., Du, L., Guo, Y., Liu, Y., & Wang, R. (2025). Evaluating large language models as raters in large-scale writing assessments: A psychometric framework for reliability and validity. Computers and Education: Artificial Intelligence, 9, 100481. https://doi.org/10.1016/j.caeai.2025.100481 [Google Scholar] [Crossref]

57. Warschauer, M., Tseng, W., Yim, S., Webster, T., Jacob, S., Du, Q., & Tate, T. (2023). The affordances and contradictions of AI-generated text for writers of English as a second or foreign language. Journal of Second Language Writing, 62, 101071. https://doi.org/10.1016/j.jslw.2023.101071 [Google Scholar] [Crossref]

58. Wei, P., Wang, X., & Dong, H. (2023). The impact of automated writing evaluation on second language writing skills of Chinese EFL learners: A randomized controlled trial. Frontiers in Psychology, 14, 1249991. https://doi.org/10.3389/fpsyg.2023.1249991 [Google Scholar] [Crossref]

59. Weigle, S. C. (2002). Assessing writing. Cambridge University Press. [Google Scholar] [Crossref]

60. Yan, X. (2025). Generative AI for the teaching, learning, and assessment of productive skills: An evidence-based approach to understanding its real impact. TESOL Quarterly. https://doi.org/10.1002/tesq.70043 [Google Scholar] [Crossref]

61. Zhai, N., & Ma, X. (2023). The effectiveness of automated writing evaluation on writing quality: A meta-analysis. Journal of Educational Computing Research, 61(4), 875–900. https://doi.org/10.1177/07356331221127300 [Google Scholar] [Crossref]

62. Zhang, Z. V., & Hyland, K. (2018). Student engagement with teacher and automated feedback on L2 writing. Assessing Writing, 36, 90–102. https://doi.org/10.1016/j.asw.2018.02.004 [Google Scholar] [Crossref]

63. Zhang, Z. V., & Hyland, K. (2022). Fostering student engagement with feedback: An integrated approach. Assessing Writing, 51, 100586. https://doi.org/10.1016/j.asw.2021.100586 [Google Scholar] [Crossref]

64. Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory into Practice, 41(2), 64–70. https://doi.org/10.1207/s15430421tip4102_2 [Google Scholar] [Crossref]

Metrics

Views & Downloads

Similar Articles