Abstract
Generative Künstliche Intelligenz (genKI) kann Lern- und Prüfungsprozesse individualisieren – je nach Einsatz kann sie dabei erweiternd oder transformierend wirken (Puentedura 2006). Zu überprüfen gilt, inwiefern solche personalisierten Prüfungen valide sind und in welchem Masse diese das Verhältnis von Standardisierung und Personalisierung in prüfungsbezogenen Prozessen verändern. Vor diesem Hintergrund untersucht die empirische Studie die Validität von zwei personalisierten Prüfungen – eine kontextualisierte, textbasierte Prüfung (erweiternd) und die Nutzung eines Prüfungsbots während der Prüfungssituation (transformierend) – anhand einer argumentativen Evidenzkette (St-Onge et al. 2017; Huggins-Manley et al. 2022), strukturiert durch die prüfungsbezogenen Prozesse scoring, administration, evaluation, decision und impact (Sinharay et al. 2025). In einem Kontrollgruppendesign (n = 18) erhielten Expert:innen u. a. aus dem Bereich Angewandte Künstliche Intelligenz entweder die von einer Lehrkraft entwickelte standardisierte Prüfung oder eine davon abgeleitete, von genKI entwickelte kontextualisierte Prüfung, die beide im Kurs «KI-Management» eines Weiterbildungsinstituts eingesetzt werden könnten. Die Ergebnisse zeigen, dass der erweiternde Personalisierungsgrad die Validität nicht signifikant beeinträchtigt oder aufwertet; ein signifikanter authentizitätsfördernder Effekt zeigt sich bei der kontextualisierten Prüfung. Für den Einsatz eines Prüfungsbots gilt: Validität ist dann gefährdet, wenn die Eigenleistung der lernenden Person nicht mehr ersichtlich ist; vorteilig wirkt der Prüfungsbot beim Geben von Feedback während und nach der Prüfung. Somit könnten Lern- und Prüfungsprozesse weiter ausdifferenziert und die Konzepte Standardisierung und Personalisierung als komplementär verstanden werden (Bates et al. 2019).
Literatur
Ali, Farhan, Doris Choy, Shanti Divaharan, Hui Yong Tay, und Wenli Chen. 2023. «Supporting self-directed learning and self-assessment using TeacherGAIA, a generative AI chatbot application: Learning approaches and prompt engineering». Learning: Research and Practice 9 (2): 135 – 47. https://doi.org/10.1080/23735082.2023.2258886.
Alier, Marc, Francisco García-Peñalvo, und Jorge D. Camba. 2024. «Generative Artificial Intelligence in Education: From Deceptive to Disruptive». International Journal of Interactive Multimedia and Artificial Intelligence 8 (März): 5. https://doi.org/10.9781/ijimai.2024.02.011.
American Psychological Association, National Council on Measurement in Education, und Joint Committee on Standards for Educational and Psychological Testing (U. S.). 1999. Standards for educational and psychological testing. Washington: American Educational Research Association.
Arslan, Burcu, Blair Lehman, Caitlin Tenison, Jesse R. Sparks, Alexis A. López, Lin Gu, und Diego Zapata-Rivera. 2024. «Opportunities and challenges of using generative AI to personalize educational assessment». Frontiers in Artificial Intelligence 7. https://doi.org/10.3389/frai.2024.1460651.
Baniasadi, Ali, Keyvan Salehi, Ebrahim Khodaie, Khosrow Bagheri Noaparast, und Balal Izanloo. 2023. «Fairness in Classroom Assessment: A Systematic Review». The Asia-Pacific Education Researcher 32 (Januar): 91 – 109. https://doi.org/10.1007/s40299-021-00636-z.
Bates, Joanna, Brett Schrewe, Rachel Ellaway, Pim Teunissen, und Chris Watling. 2019. «Embracing standardisation and contextualisation in medical education». Medical Education 53 (1). https://doi.org/10.1111/medu.13740.
Bearman, Margaret, und Rosemary Luckin. 2020. «Preparing University Assessment for a World with AI: Tasks for Human Intelligence». In Re-imagining University Assessment in a Digital World, herausgegeben von Margaret Bearman, Phillip Dawson, Rola Ajjawi, Joanna Tai, und David Boud, 49 – 63. Cham: Springer. https://doi.org/10.1007/978-3-030-41956-1_5.
Bennett, Randy E. 2024. «Personalizing Assessment: Dream or Nightmare?» Educational Measurement: Issues and Practice n/a (n/a). https://doi.org/10.1111/emip.12652.
Bortz, Jürgen, und Nicola Döring. 2006. Forschungsmethoden und Evaluation für Human- und Sozialwissenschaftler (Sonderausgabe der 4. Auflage). Berlin, Heidelberg: Springer. https://doi.org/10.1007/978-3-540-33306-7.
Boud, David. 2000. «Sustainable Assessment: Rethinking assessment for the learning society». Studies in Continuing Education 22 (2): 151 – 67. https://doi.org/10.1080/713695728.
Boud, David, und Rebeca Soler. 2016. «Sustainable assessment revisited». Assessment & Evaluation in Higher Education 41 (3): 400 – 13. https://doi.org/10.1080/02602938.2015.1018133.
Buzick, Heather M., Jodi M. Casabianca, und Melissa L. Gholson. 2023. «Personalizing Large-Scale Assessment in Practice». Educational Measurement: Issues and Practice 42 (2): 5 – 11. https://doi.org/10.1111/emip.12551.
Chan, Kuang Wen, Farhan Ali, Joonhyeong Park, Kah Shen Brandon Sham, Erdalyn Yeh Thong Tan, Francis Woon Chien Chong, Kun Qian, und Guan Kheng Sze. 2025. «Automatic item generation in various STEM subjects using large language model prompting». Computers and Education: Artificial Intelligence 8 (Juni): 100344. https://doi.org/10.1016/j.caeai.2024.100344.
Chang, Jina, Joonhyeong Park, und Jisun Park. 2023. «Using an Artificial Intelligence Chatbot in Scientific Inquiry: Focusing on a Guided-Inquiry Activity Using Inquirybot». Asia-Pacific Science Education (Leiden, Niederlande) 9 (1): 44 – 74. https://doi.org/10.1163/23641177-bja10062.
Circi, Ruhan, Juanita Hicks, und Emmanuel Sikali. 2023. «Automatic item generation: foundations and machine learning-based approaches for assessments». Frontiers in Education 8. https://doi.org/10.3389/feduc.2023.858273.
Cohen, Jacob. 1988. Statistical Power Analysis for the Behavioral Sciences. 2nd Edition. New York: Routledge.
Corbin, Thomas, Phillip Dawson, und Danny Liu. 2025. «Talk is cheap: why structural assessment changes are needed for a time of GenAI». Assessment & Evaluation in Higher Education Mai 15: 1 – 11. https://doi.org/10.1080/02602938.2025.2503964.
Crooks, Terry, Michael Kane, und Allan Cohen. 1996. «Threats to the Valid Use of Assessments». Assessment in Education: Principles, Policy & Practice 3 (November): 265 – 86. https://doi.org/10.1080/0969594960030302.
Dawson, Phillip, Margaret Bearman, Mollie Dollinger, und David Boud. 2024. «Validity matters more than cheating». Assessment & Evaluation in Higher Education 49 (7): 1005 – 16. https://doi.org/10.1080/02602938.2024.2386662.
Furze, Leon, Mike Perkins, Jasper Roe, und Jason MacVaugh. 2024. «The AI Assessment Scale (AIAS) in action: A pilot implementation of GenAI-supported assessment». Australasian Journal of Educational Technology, Online-Vorab-Publikation, Oktober 16. https://doi.org/10.14742/ajet.9434.
Hedges, Larry, und Ingram Olkin. 1985. «Statistical Methods in Meta-Analysis». Orlando: Academic Press. https://doi.org/10.1016/C2009-0-03396-0.
Huggins-Manley, A. Corinne, Brandon M. Booth, und Sidney K. DʼMello. 2022. «Toward Argument-Based Fairness with an Application to AI-Enhanced Educational Assessments». Journal of Educational Measurement 59 (3): 362 – 88. https://doi.org/10.1111/jedm.12334.
Kaldaras, Leonora, Hope O. Akaeze, und Mark D. Reckase. 2024. «Developing valid assessments in the era of generative artificial intelligence». Frontiers in Education 9 . https://doi.org/10.3389/feduc.2024.1399377.
Klar, Maria, und Johannes Schleiss. 2024. «Künstliche Intelligenz im Kontext von Kompetenzen, Prüfungen und Lehr-Lern-Methoden: Alte und neue Gestaltungsfragen». JFMH23: Spannungsfeld Der Digitalen Kompetenz. MedienPädagogik: Zeitschrift für Theorie und Praxis der Medienbildung 58 (JFMH2023): 41 – 57. https://doi.org/10.21240/mpaed/58/2024.03.24.X.
Labadze, Lasha, Maya Grigolia, und Lela Machaidze. 2023. «Role of AI chatbots in education: systematic literature review». International Journal of Educational Technology in Higher Education 20 (1): 56. https://doi.org/10.1186/s41239-023-00426-1.
Lakens, Daniel. 2013. «Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs». Frontiers in Psychology 4. https://doi.org/10.3389/fpsyg.2013.00863.
Long, Duri, und Brian Magerko. 2020. «What is AI Literacy? Competencies and Design Considerations». Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (New York, NY, USA), CHI ʼ20, 1 – 16. https://doi.org/10.1145/3313831.3376727.
Luo (Jess), Jiahui. 2024. «A critical review of GenAI policies in higher education assessment: a call to reconsider the ‹originality› of studentsʼ work». Assessment & Evaluation in Higher Education 49 (5): 651 – 64. https://doi.org/10.1080/02602938.2024.2309963.
Mann, Henry, und Donald Whitney. 1947. «On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other». The Annals of Mathematical Statistics 18 (1): 50 – 60. https://doi.org/10.1214/aoms/1177730491.
Murillo, F. Javier, und Nina Hidalgo. 2020. «Fair student assessment: A phenomenographic study on teachersʼ conceptions». Studies in Educational Evaluation 65 (Juni): 100860. https://doi.org/10.1016/j.stueduc.2020.100860.
Nicola-Richmond, Kelli, Phillip Dawson, Helen Partridge, und Susie Macfarlane. 2025. «It takes a village… Program-wide approaches to redesigning assessment in a time of generative artificial intelligence (GenAI)». Curriculum and Assessment Design. Journal of University Teaching and Learning Practice. https://doi.org/10.53761/zpp2ja61.
Nieminen, Juuso Henrik, Mollie Dollinger, und Rachel Finneran. 2025. «‹There was very little room for me to be me›: the lived tensions between assessment standardisation and student diversity». Assessment & Evaluation in Higher Education 50 (2): 308 – 22. https://doi.org/10.1080/02602938.2024.2388699.
Puentedura, Ruben. 2006. «Transformation, Technology, and Education». http://hippasus.com/resources/tte/puentedura_tte.pdf.
Reinhold, Lhea, und Marion Händel. 2025. «Bewertung durch eine künstliche Intelligenz? Auswertungs- und Interpretationsobjektivität von ChatGPT-4o bei der Bewertung von Lerntagebucheinträgen». MEDIDA24: 1. Tagungsband des AK Mediendidaktik der DGfE-Sektion Medienpädagogik. MedienPädagogik: Zeitschrift für Theorie und Praxis der Medienbildung 65 (MEDIDA24): 227 – 50. https://doi.org/10.21240/mpaed/65/2025.08.03.X.
Reinmann, Gabi. 2021. «Prüfungstypen, -formate, -formen oder -szenarien?» Impact Free 36. https://gabi-reinmann.de/wp-content/uploads/2021/06/Impact_Free_36.pdf.
Santoso, Agung, Alice Whita Savira, und Robertus Landung Eko Prihatmoko. 2022. «Validity evidence based on content: Controversies and quantification». 5 – 16. https://doi.org/10.37517/978-1-74286-697-0-01.
Saxena, Ritcha, Kevin Carnevale, Oleg Yakymovych, Michael Salzle, Kapil Sharma, und Ritwik Raj Saxena. 2023. «Precision, Personalization, and Progress: Traditional and Adaptive Assessment in Undergraduate Medical Education». Articles. Innovative Research Thoughts 9 (4): 216 – 23. https://doi.org/10.36676/irt.2023-v9i4-029.
Schaper, Niclas. 2021. «Prüfen in der Hochschullehre». In Handbuch Hochschuldidaktik, herausgegeben von Robert Kordts-Freudinger, Niclas Schaper, Antonia Scholkmann, und Birgit Szczyrba, 87 – 102. Stuttgart: wbv Publikation. https://www.utb.de/doi/full/10.36198/9783838554082-87-102.
Sinharay, Sandip, Randy E. Bennett, Michael Kane, und Jesse R. Sparks. 2025. «Validation for Personalized Assessments: A Threats-to-Validity Approach». Journal of Educational Measurement n/a (n/a). https://doi.org/10.1111/jedm.12434.
Sireci, Stephen G. 2020. «Standardization and UNDERSTANDardization in Educational Assessment». Educational Measurement: Issues and Practice 39 (3): 100 – 5. https://doi.org/10.1111/emip.12377.
St-Onge, Christina, Meredith Young, Kevin W. Eva, und Brian Hodges. 2017. «Validity: One Word with a Plurality of Meanings». Advances in Health Sciences Education: Theory and Practice (Netherlands) 22 (4): 853 – 67. https://doi.org/10.1007/s10459-016-9716-3.
Tierney, Robin D. 2014. «Fairness as a multifaceted quality in classroom assessment». Studies in Educational Evaluation 43 (Dezember): 55 – 69. https://doi.org/10.1016/j.stueduc.2013.12.003.
Walker, Michael E., Margarita Olivera-Aguilar, Blair Lehman, Cara Laitusis, Danielle Guz-man-Orth, und Melissa Gholson. 2023. «Culturally Responsive Assessment: Provisional Principles». ETS Research Report Series 2023 (1): 1 – 24. https://doi.org/10.1002/ets2.12374.
Wilcoxon, Frank. 1945. «Individual Comparisons by Ranking Methods». Biometrics Bulletin 1 (6): 80 – 8. JSTOR. https://doi.org/10.2307/3001968.
Wollny, Sebastian, Jan Schneider, Daniele Di Mitri, Joshua Weidlich, Marc Rittberger, und Hendrik Drachsler. 2021. «Are We There Yet? – A Systematic Literature Review on Chatbots in Education». Frontiers in Artificial Intelligence 4 . https://doi.org/10.3389/frai.2021.654924.
