Abstract
Generative Artificial Intelligence (genAI) has the potential to individualize learning and assessment processes – depending on its mode of application, genAI can have either an augmentative or a transformative effect (Puentedura 2006). A central question is to what extent such personalized assessments can be considered valid, and to what degree they alter the relationship between standardization and personalization within assessment-related processes. Against this backdrop, the empirical study investigates the validity of two forms of personalized assessment – one contextualized, text-based assessment (augmentative) and the use of an assessment bot during the assessment situation (transformative) – through an argumentative chain of evidence (St-Onge et al. 2017; Huggins-Manley et al. 2022) structured according to the assessment processes of scoring, administration, evaluation, decision, and impact (Sinharay et al. 2025). Using a control group design (n = 18), experts, including specialists in Applied Artificial Intelligence, received either a teacher-developed standardized assessment or a genAI-developed contextualized assessment derived from it, both of which could be implemented in the course «AI Management» at a continuing education institute. The findings indicate that the augmentative level of personalization neither significantly impairs nor enhances validity; however, a significant effect promoting authenticity was observed for the contextualized assessment. Regarding the use of an assessment bot, validity appears compromised when the learnerʼs individual contribution becomes indistinguishable, yet the bot proves advantageous for providing feedback during and after the assessment. Consequently, learning and assessment processes could be further differentiated, and the concepts of standardization and personalization could be understood to be complementary (Bates et al. 2019).
References
Ali, Farhan, Doris Choy, Shanti Divaharan, Hui Yong Tay, und Wenli Chen. 2023. «Supporting self-directed learning and self-assessment using TeacherGAIA, a generative AI chatbot application: Learning approaches and prompt engineering». Learning: Research and Practice 9 (2): 135 – 47. https://doi.org/10.1080/23735082.2023.2258886.
Alier, Marc, Francisco García-Peñalvo, und Jorge D. Camba. 2024. «Generative Artificial Intelligence in Education: From Deceptive to Disruptive». International Journal of Interactive Multimedia and Artificial Intelligence 8 (März): 5. https://doi.org/10.9781/ijimai.2024.02.011.
American Psychological Association, National Council on Measurement in Education, und Joint Committee on Standards for Educational and Psychological Testing (U. S.). 1999. Standards for educational and psychological testing. Washington: American Educational Research Association.
Arslan, Burcu, Blair Lehman, Caitlin Tenison, Jesse R. Sparks, Alexis A. López, Lin Gu, und Diego Zapata-Rivera. 2024. «Opportunities and challenges of using generative AI to personalize educational assessment». Frontiers in Artificial Intelligence 7. https://doi.org/10.3389/frai.2024.1460651.
Baniasadi, Ali, Keyvan Salehi, Ebrahim Khodaie, Khosrow Bagheri Noaparast, und Balal Izanloo. 2023. «Fairness in Classroom Assessment: A Systematic Review». The Asia-Pacific Education Researcher 32 (Januar): 91 – 109. https://doi.org/10.1007/s40299-021-00636-z.
Bates, Joanna, Brett Schrewe, Rachel Ellaway, Pim Teunissen, und Chris Watling. 2019. «Embracing standardisation and contextualisation in medical education». Medical Education 53 (1). https://doi.org/10.1111/medu.13740.
Bearman, Margaret, und Rosemary Luckin. 2020. «Preparing University Assessment for a World with AI: Tasks for Human Intelligence». In Re-imagining University Assessment in a Digital World, herausgegeben von Margaret Bearman, Phillip Dawson, Rola Ajjawi, Joanna Tai, und David Boud, 49 – 63. Cham: Springer. https://doi.org/10.1007/978-3-030-41956-1_5.
Bennett, Randy E. 2024. «Personalizing Assessment: Dream or Nightmare?» Educational Measurement: Issues and Practice n/a (n/a). https://doi.org/10.1111/emip.12652.
Bortz, Jürgen, und Nicola Döring. 2006. Forschungsmethoden und Evaluation für Human- und Sozialwissenschaftler (Sonderausgabe der 4. Auflage). Berlin, Heidelberg: Springer. https://doi.org/10.1007/978-3-540-33306-7.
Boud, David. 2000. «Sustainable Assessment: Rethinking assessment for the learning society». Studies in Continuing Education 22 (2): 151 – 67. https://doi.org/10.1080/713695728.
Boud, David, und Rebeca Soler. 2016. «Sustainable assessment revisited». Assessment & Evaluation in Higher Education 41 (3): 400 – 13. https://doi.org/10.1080/02602938.2015.1018133.
Buzick, Heather M., Jodi M. Casabianca, und Melissa L. Gholson. 2023. «Personalizing Large-Scale Assessment in Practice». Educational Measurement: Issues and Practice 42 (2): 5 – 11. https://doi.org/10.1111/emip.12551.
Chan, Kuang Wen, Farhan Ali, Joonhyeong Park, Kah Shen Brandon Sham, Erdalyn Yeh Thong Tan, Francis Woon Chien Chong, Kun Qian, und Guan Kheng Sze. 2025. «Automatic item generation in various STEM subjects using large language model prompting». Computers and Education: Artificial Intelligence 8 (Juni): 100344. https://doi.org/10.1016/j.caeai.2024.100344.
Chang, Jina, Joonhyeong Park, und Jisun Park. 2023. «Using an Artificial Intelligence Chatbot in Scientific Inquiry: Focusing on a Guided-Inquiry Activity Using Inquirybot». Asia-Pacific Science Education (Leiden, Niederlande) 9 (1): 44 – 74. https://doi.org/10.1163/23641177-bja10062.
Circi, Ruhan, Juanita Hicks, und Emmanuel Sikali. 2023. «Automatic item generation: foundations and machine learning-based approaches for assessments». Frontiers in Education 8. https://doi.org/10.3389/feduc.2023.858273.
Cohen, Jacob. 1988. Statistical Power Analysis for the Behavioral Sciences. 2nd Edition. New York: Routledge.
Corbin, Thomas, Phillip Dawson, und Danny Liu. 2025. «Talk is cheap: why structural assessment changes are needed for a time of GenAI». Assessment & Evaluation in Higher Education Mai 15: 1 – 11. https://doi.org/10.1080/02602938.2025.2503964.
Crooks, Terry, Michael Kane, und Allan Cohen. 1996. «Threats to the Valid Use of Assessments». Assessment in Education: Principles, Policy & Practice 3 (November): 265 – 86. https://doi.org/10.1080/0969594960030302.
Dawson, Phillip, Margaret Bearman, Mollie Dollinger, und David Boud. 2024. «Validity matters more than cheating». Assessment & Evaluation in Higher Education 49 (7): 1005 – 16. https://doi.org/10.1080/02602938.2024.2386662.
Furze, Leon, Mike Perkins, Jasper Roe, und Jason MacVaugh. 2024. «The AI Assessment Scale (AIAS) in action: A pilot implementation of GenAI-supported assessment». Australasian Journal of Educational Technology, Online-Vorab-Publikation, Oktober 16. https://doi.org/10.14742/ajet.9434.
Hedges, Larry, und Ingram Olkin. 1985. «Statistical Methods in Meta-Analysis». Orlando: Academic Press. https://doi.org/10.1016/C2009-0-03396-0.
Huggins-Manley, A. Corinne, Brandon M. Booth, und Sidney K. DʼMello. 2022. «Toward Argument-Based Fairness with an Application to AI-Enhanced Educational Assessments». Journal of Educational Measurement 59 (3): 362 – 88. https://doi.org/10.1111/jedm.12334.
Kaldaras, Leonora, Hope O. Akaeze, und Mark D. Reckase. 2024. «Developing valid assessments in the era of generative artificial intelligence». Frontiers in Education 9 . https://doi.org/10.3389/feduc.2024.1399377.
Klar, Maria, und Johannes Schleiss. 2024. «Künstliche Intelligenz im Kontext von Kompetenzen, Prüfungen und Lehr-Lern-Methoden: Alte und neue Gestaltungsfragen». JFMH23: Spannungsfeld Der Digitalen Kompetenz. MedienPädagogik: Zeitschrift für Theorie und Praxis der Medienbildung 58 (JFMH2023): 41 – 57. https://doi.org/10.21240/mpaed/58/2024.03.24.X.
Labadze, Lasha, Maya Grigolia, und Lela Machaidze. 2023. «Role of AI chatbots in education: systematic literature review». International Journal of Educational Technology in Higher Education 20 (1): 56. https://doi.org/10.1186/s41239-023-00426-1.
Lakens, Daniel. 2013. «Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs». Frontiers in Psychology 4. https://doi.org/10.3389/fpsyg.2013.00863.
Long, Duri, und Brian Magerko. 2020. «What is AI Literacy? Competencies and Design Considerations». Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (New York, NY, USA), CHI ʼ20, 1 – 16. https://doi.org/10.1145/3313831.3376727.
Luo (Jess), Jiahui. 2024. «A critical review of GenAI policies in higher education assessment: a call to reconsider the ‹originality› of studentsʼ work». Assessment & Evaluation in Higher Education 49 (5): 651 – 64. https://doi.org/10.1080/02602938.2024.2309963.
Mann, Henry, und Donald Whitney. 1947. «On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other». The Annals of Mathematical Statistics 18 (1): 50 – 60. https://doi.org/10.1214/aoms/1177730491.
Murillo, F. Javier, und Nina Hidalgo. 2020. «Fair student assessment: A phenomenographic study on teachersʼ conceptions». Studies in Educational Evaluation 65 (Juni): 100860. https://doi.org/10.1016/j.stueduc.2020.100860.
Nicola-Richmond, Kelli, Phillip Dawson, Helen Partridge, und Susie Macfarlane. 2025. «It takes a village… Program-wide approaches to redesigning assessment in a time of generative artificial intelligence (GenAI)». Curriculum and Assessment Design. Journal of University Teaching and Learning Practice. https://doi.org/10.53761/zpp2ja61.
Nieminen, Juuso Henrik, Mollie Dollinger, und Rachel Finneran. 2025. «‹There was very little room for me to be me›: the lived tensions between assessment standardisation and student diversity». Assessment & Evaluation in Higher Education 50 (2): 308 – 22. https://doi.org/10.1080/02602938.2024.2388699.
Puentedura, Ruben. 2006. «Transformation, Technology, and Education». http://hippasus.com/resources/tte/puentedura_tte.pdf.
Reinhold, Lhea, und Marion Händel. 2025. «Bewertung durch eine künstliche Intelligenz? Auswertungs- und Interpretationsobjektivität von ChatGPT-4o bei der Bewertung von Lerntagebucheinträgen». MEDIDA24: 1. Tagungsband des AK Mediendidaktik der DGfE-Sektion Medienpädagogik. MedienPädagogik: Zeitschrift für Theorie und Praxis der Medienbildung 65 (MEDIDA24): 227 – 50. https://doi.org/10.21240/mpaed/65/2025.08.03.X.
Reinmann, Gabi. 2021. «Prüfungstypen, -formate, -formen oder -szenarien?» Impact Free 36. https://gabi-reinmann.de/wp-content/uploads/2021/06/Impact_Free_36.pdf.
Santoso, Agung, Alice Whita Savira, und Robertus Landung Eko Prihatmoko. 2022. «Validity evidence based on content: Controversies and quantification». 5 – 16. https://doi.org/10.37517/978-1-74286-697-0-01.
Saxena, Ritcha, Kevin Carnevale, Oleg Yakymovych, Michael Salzle, Kapil Sharma, und Ritwik Raj Saxena. 2023. «Precision, Personalization, and Progress: Traditional and Adaptive Assessment in Undergraduate Medical Education». Articles. Innovative Research Thoughts 9 (4): 216 – 23. https://doi.org/10.36676/irt.2023-v9i4-029.
Schaper, Niclas. 2021. «Prüfen in der Hochschullehre». In Handbuch Hochschuldidaktik, herausgegeben von Robert Kordts-Freudinger, Niclas Schaper, Antonia Scholkmann, und Birgit Szczyrba, 87 – 102. Stuttgart: wbv Publikation. https://www.utb.de/doi/full/10.36198/9783838554082-87-102.
Sinharay, Sandip, Randy E. Bennett, Michael Kane, und Jesse R. Sparks. 2025. «Validation for Personalized Assessments: A Threats-to-Validity Approach». Journal of Educational Measurement n/a (n/a). https://doi.org/10.1111/jedm.12434.
Sireci, Stephen G. 2020. «Standardization and UNDERSTANDardization in Educational Assessment». Educational Measurement: Issues and Practice 39 (3): 100 – 5. https://doi.org/10.1111/emip.12377.
St-Onge, Christina, Meredith Young, Kevin W. Eva, und Brian Hodges. 2017. «Validity: One Word with a Plurality of Meanings». Advances in Health Sciences Education: Theory and Practice (Netherlands) 22 (4): 853 – 67. https://doi.org/10.1007/s10459-016-9716-3.
Tierney, Robin D. 2014. «Fairness as a multifaceted quality in classroom assessment». Studies in Educational Evaluation 43 (Dezember): 55 – 69. https://doi.org/10.1016/j.stueduc.2013.12.003.
Walker, Michael E., Margarita Olivera-Aguilar, Blair Lehman, Cara Laitusis, Danielle Guz-man-Orth, und Melissa Gholson. 2023. «Culturally Responsive Assessment: Provisional Principles». ETS Research Report Series 2023 (1): 1 – 24. https://doi.org/10.1002/ets2.12374.
Wilcoxon, Frank. 1945. «Individual Comparisons by Ranking Methods». Biometrics Bulletin 1 (6): 80 – 8. JSTOR. https://doi.org/10.2307/3001968.
Wollny, Sebastian, Jan Schneider, Daniele Di Mitri, Joshua Weidlich, Marc Rittberger, und Hendrik Drachsler. 2021. «Are We There Yet? – A Systematic Literature Review on Chatbots in Education». Frontiers in Artificial Intelligence 4 . https://doi.org/10.3389/frai.2021.654924.
