Abstract
Künstliche Intelligenz (KI) kann im Prozess der Leistungsbewertung assistieren und diesen transformieren. Besonders lohnend scheint eine KI-Assistenz bei der Bewertung von komplexem, geschriebenem Text. Jedoch ist der Einsatz von KI im Bewertungsprozess «hochriskant» (EU 2024) und bedarf umfangreicher Analysen. Die vorliegende Studie untersucht, inwiefern ChatGPT-4o die Auswertung und Interpretation von Lerntagebucheinträgen objektiv vornehmen kann. Dafür werden 757 Lerntagebucheinträge aus der geförderten Weiterbildung in Deutschland von Mensch und Maschine bewertet. Sowohl Mensch als auch Maschine erhalten hierzu Kriterien, nach denen die Bewertung vorzunehmen ist; ChatGPT-4o wird diesbezüglich mit einem Prompt unterstützt. Die Übereinstimmung der Bewertungen wird anhand der Masse Sensitivität und Spezifität gemessen. Die Ergebnisse zeigen, dass die Bewertungsvorschläge von ChatGPT-4o eine moderate bis hohe Übereinstimmung mit den menschlichen Bewertungen aufweisen; gleichzeitig neigt ChatGPT-4o jedoch zu einer optimistischen Bewertung der Lerntagebucheinträge. Die Ergebnisse weisen darauf hin, dass eine hybride Intelligenz, also eine Kombination der Stärken von Mensch und Maschine, gewinnbringend für Bewertungsprozesse sein kann. Künftig denkbar sind halbautomatisierte Bewertungsprozesse von Lerntagebucheinträgen, in denen die KI die Bewertung der Lerntagebucheinträge übernimmt und Lehrkräfte bei kritischen Fällen regulierend eingreifen. So könnte die Korrektureffizienz ohne bedeutende Qualitätsverluste gesteigert werden.
Literatur
Alers, Hani, Aleksandra Malinowska, Gregory Meghoe, und Enso Apfel. 2024. «Using ChatGPT-4 to Grade Open Question Exams». In Advances in Information and Communication Proceedings of the 2024 Future of Information and Communication Conference (FICC), Volume 1, 919: 1–9. Lecture Notes in Networks and Systems. Berlin: Springer Nature. https://doi.org/10.1007/978-3-031-53960-2_1.
Benischek, Isabella, Angela Forstner-Ebhart, Hubert Schaupp, und Herbert Schwetz, Hrsg. 2012. «Empirische Forschung zu schulischen Handlungsfeldern. Ergebnisse der ARGE Bildungsforschung an pädagogischen Hochschulen in Österreich. 2.» Austria: Forschung und Wissenschaft. Erziehungswissenschaft. 15. Berlin u. a.: Lit.
Bernius, Jan Philip, Stephan Krusche, und Bernd Bruegge. 2022. «Machine learning based feedback on textual student answers in large courses». Computers and Education: Artificial Intelligence 3 (Januar): 100081. https://doi.org/10.1016/j.caeai.2022.100081.
Birkel, Peter, und Claudia Birkel. 2002. «Wie einig sind sich Lehrer bei der Aufsatz beurteilung ? Eine Replikationsstudie zur Untersuchung von Rudolf Weiss.» Psychologie in Erziehung und Unterricht 49 (3): 219–24.
Brookhart, Susan M. 2018. «Appropriate Criteria: Key to Effective Rubrics». Frontiers in Education 3. https://doi.org/10.3389/feduc.2018.00022.
Brookhart, Susan M., Thomas R. Guskey, Alex J. Bowers, James H. McMillan, Jeffrey K. Smith, Lisa F. Smith, Michael T. Stevens, und Megan E. Welsh. 2016. «A Century of Grading Research: Meaning and Value in the Most Common Educational Measure». Review of Educational Research 86 (4): 803–48. https://doi.org/10.3102/0034654316672069.
Brüggemann, Tim, und Claudia Wiepcke. 2023. «Der EdTech-Index (ETX)», herausgegeben von Claudia Wiepcke. https://nbn-resolving.org/urn:nbn:de:bsz:751-opus4-4570.
Bürgermeister, Anika, Inga Glogger-Frey, und Henrik Saalbach. 2021. «Supporting Peer Feedback on Learning Strategies: Effects on Self-Efficacy and Feedback Quality». Psychology Learning & Teaching 20 (3): 383–404. https://doi.org/10.1177/14757257211016604.
Celik, Ismail. 2023. «Towards Intelligent-TPACK: An empirical study on teachers’ professional knowledge to ethically integrate artificial intelligence (AI)-based tools into education». Computers in Human Behavior 138 (Januar):107468. https://doi.org/10.1016/j.chb.2022.107468.
Dellermann, Dominik, Philipp Ebel, Matthias Söllner, und Jan Marco Leimeister. 2019. «Hybrid Intelligence». Business & Information Systems Engineering,: Vol. 61, No. 5, S.637–43. http://dx.doi.org/10.1007/s12599-019-00595-2.
Ehlers, Jan P., Christian Guetl, Susan Höntzsch, Claus A. Usener, und Susanne Gruttmann. 2013. «Prüfen mit Computer und Internet. Didaktik, Methodik und Organisation von E-Assessment». In Lehrbuch für Lernen und Lehren mit Technologien, herausgegeben von Martin Ebner und Sandra Schön, 2. Auflage. Berlin: epubli GmbH: 227–237.
Elsäßer, Sibylle Dorothee. 2017. «Komponenten von schulischen Leistungen: Eine Analyse zu Einflussfaktoren auf die Notengebung in der Grundschule». München: Ludwig-Maximilians-Universität, Elektronische Hochschulschriften. http://nbn-resolving.de/urn:nbn:de:bvb:19-226439.
EU. 2024. Verordnung (EU) 2024/1689 des Europäischen Parlaments und des Rates vom 16. Juli 2024. https://eur-lex.europa.eu/legal-content/DE/TXT/?uri=CELEX%3A32024R1689.
Gentile, Manuel, Giuseppe Città, Salvatore Perna, und Mario Allegra. 2023. «Do we still need teachers? Navigating the paradigm shift of the teacher’s role in the AI era». Frontiers in Education 8. https://doi.org/10.3389/feduc.2023.1161777.
Glogger, Inga, Rolf Schwonke, Lars Holzäpfel, Matthias Nückles, und Alexander Renkl. 2012. «Learning Strategies Assessed by Journal Writing: Prediction of Learning Outcomes by Quantity, Quality, and Combinations of Learning Strategies». Journal of Educational Psychology 104 (Januar):452–68. https://doi.org/10.1037/a0026683.
Goh, Ethan, Robert Gallo, Jason Hom, Eric Strong, Yingjie Weng, Hannah Kerman, Joséphine A. Cool, et al. 2024. «Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial». JAMA Network Open 7 (10): e2440969–e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969.
González-Calatayud, Víctor, Paz Prendes-Espinosa, und Rosabel Roig-Vila. 2021. «Artificial Intelligence for Student Assessment: A Systematic Review». Applied Sciences 11 (12). https://doi.org/10.3390/app11125467.
Grienberger, Katharina, Britta Matthes, und Wiebke Paulus. 2024. «Folgen des technologischen Wandels für den Arbeitsmarkt: Vor allem Hochqualifizierte bekommen die Digitalisierung verstärkt zu spüren». IAB-Kurzbericht 2024, 5. Nürnberg: Institut für Arbeitsmarkt- und Berufsforschung. https://doi.org/10.48720/IAB.KB.2405.
Hattie, John, Deb Masters, und Kate Birch. 2015. Visible Learning into Action International Case Studies of Impact. London: Routledge. https://doi.org/10.4324/9781315722603.
Hussein, Mohamed Abdellatif, Hesham Hassan, und Mohammad Nassef. 2019. «Automated language essay scoring systems: a literature review». PeerJ Computer Science 5:e208. https://doi.org/10.7717/peerj-cs.208.
Ingenkamp, Karlheinz, und Urban Lissmann. 2008. Lehrbuch der pädagogischen Diagnostik. 6., neu ausgestattete Aufl. Beltz Pädagogik. Weinheim u. a.: Beltz.
Jacobsen, Lucas, und Kira Weber. 2023. «The Promises and Pitfalls of ChatGPT as a Feedback Provider in Higher Education: An Exploratory Study of Prompt Engineering and the Quality of AI-Driven Feedback.» https://doi.org/10.31219/osf.io/cr257.
Jönsson, Anders, Andreia Balan, und Eva Hartell. 2021. «Analytic or holistic? A study about how to increase the agreement in teachers’ grading». Assessment in Education: Principles, Policy & Practice 28 (3): 212–27. https://doi.org/10.1080/0969594X.2021.1884041.
Jukiewicz, Marcin. 2024. «The future of grading programming assignments in education: The role of ChatGPT in automating the assessment and feedback process». Thinking Skills and Creativity 52:101522. https://doi.org/10.1016/j.tsc.2024.101522.
Kircher, Ernst, Raimund Girwidz, und Hans E. Fischer. 2020. Physikdidaktik: Grundlagen. 4. Auflage. Berlin: Springer. https://doi.org/10.1007/978-3-662-59490-2.
Kooli, Chokri, und Nadia Yusuf. 2024. «Transforming Educational Assessment: Insights Into the Use of ChatGPT and Large Language Models in Grading». International Journal of Human–Computer Interaction 0 (0): 1–12. https://doi.org/10.1080/10447318.2024.2338330.
Lameras, Petros, und Sylvester Arnab. 2022. «Power to the Teachers: An Exploratory Review on Artificial Intelligence in Education». Information 13 (1). https://doi.org/10.3390/info13010014.
Lerche, Thomas. 2022. Leistungsbeurteilung an Schulen. München: Lehrstuhl für Schulpädagogik, Ludwig-Maximilians-Universität München. https://epub.ub.uni-muenchen.de/91872/.
Li, Shengjie, und Vincent Ng. 2024. «Automated Essay Scoring: Recent Successes and Future Directions». In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, herausgegeben von Kate Larson, 8114–22. International Joint Conferences on Artificial Intelligence Organization. https://doi.org/10.24963/ijcai.2024/897.
Lundgren, Magnus. 2024. «Large Language Models in Student Assessment: Comparing ChatGPT and Human Graders.» https://doi.org/10.13140/RG.2.2.27630.42561.
Mishra, Punya, Melissa Warr, und Rezwana Islam. 2023. «TPACK in the age of ChatGPT and Generative AI». Journal of Digital Learning in Teacher Education 39 (4): 235–51. https://doi.org/10.1080/21532974.2023.2247480.
Mizumoto, Atsushi, und Masaki Eguchi. 2023. «Exploring the potential of using an AI language model for automated essay scoring». Research Methods in Applied Linguistics 2 (2): 100050. https://doi.org/10.1016/j.rmal.2023.100050.
Molenaar, Inge. 2021. «Personalisation of learning: Towards hybrid human-AI learning technologies». In OECD digital education outlook 2021: Pushing the frontiers with AI, blockchain, and robots, herausgegeben von S. Vincent-Lancrin, 57–77. OECD. https://doi.org/10.1787/589b283f-en.
Molenaar, Inge. 2022. «Towards Hybrid human‐AI Learning Technologies». European Journal of Education 57 (4): 632–45. https://doi.org/10.1111/ejed.12527.
Möller, Jens, Thorben Jansen, Johanna Fleckenstein, Nils Machts, Jennifer Meyer, und Raja Reble. 2022. «Judgment accuracy of German student texts: Do teacher experience and content knowledge matter?» Teaching and Teacher Education 119 (November):103879. https://doi.org/10.1016/j.tate.2022.103879.
Naujoks, Nick, und Marion Händel. 2020. «Nur vertiefen oder auch wiederholen? Differenzielle Verläufe kognitiver Lernstrategien im Semester». Unterrichtswissenschaft 48 (2): 221–41. https://doi.org/10.1007/s42010-019-00062-7.
Ninaus, Manuel, und Michael Sailer. 2022. «Closing the loop – The human role in artificial intelligence for education». Frontiers in Psychology 13. https://doi.org/10.3389/fpsyg.2022.956798.
Ning, Yimin, Cheng Zhang, Binyan Xu, Ying Zhou, und Tommy Tanu Wijaya. 2024. «Teachers’ AI-TPACK: Exploring the Relationship between Knowledge Elements». Sustainability 16 (3). https://doi.org/10.3390/su16030978.
Nückles, Matthias, Sandra Hübner, Sandra Dümer, und Alexander Renkl. 2010. «Expertise reversal effects in writing-to-learn». Instructional Science 38 (Mai):237–58. https://doi.org/10.1007/s11251-009-9106-9.
Nückles, Matthias, Julian Roelle, Inga Glogger-Frey, Julia Waldeyer, und Alexander Renkl. 2020. «The Self-Regulation-View in Writing-to-Learn: Using Journal Writing to Optimize Cognitive Load in Self-Regulated Learning». Educational Psychology Review 32 (4): 1089–1126. https://doi.org/10.1007/s10648-020-09541-1.
Pinto, Gustavo, Isadora Cardoso-Pereira, Danilo Monteiro, Danilo Lucena, Alberto Souza, und Kiev Gama. 2023. «Large Language Models for Education: Grading Open-Ended Questions Using ChatGPT». In Proceedings of the XXXVII Brazilian Symposium on Software Engineering, 293–302. SBES ’23. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3613372.3614197.
Ramesh, Dadi, und Suresh Kumar Sanampudi. 2022. «An automated essay scoring systems: a systematic literature review». Artificial Intelligence Review 55 (3): 2495–2527. https://doi.org/10.1007/s10462-021-10068-2.
Reddy, Y., und Heidi Andrade. 2010. «A review of rubric use in higher education». Assessment & Evaluation in Higher Education – ASSESS EVAL HIGH EDUC 35 (Juli): 435–48. https://doi.org/10.1080/02602930902862859.
Schraw, Gregory, Fred Kuch, und Antonio P. Gutierrez. 2013. «Measure for measure: Calibrating ten commonly used calibration scores». Calibrating Calibration: Creating Conceptual Clarity to Guide Measurement and Calculation 24 (April):48–57. https://doi.org/10.1016/j.learninstruc.2012.08.007.
Sonnenschein, Katharina. 2015. Die Objektivität der Leistungsbewertung: Inwieweit sind Leistungsbeurteilungen von Lehrkräften einheitlich und somit vergleichbar? Hamburg: Diplomica.
Spörer, Nadine, und Joachim C. Brunstein. 2006. «Erfassung selbstregulierten Lernens mit Selbstberichtsverfahren». Zeitschrift für Pädagogische Psychologie 20 (3): 147–60. https://doi.org/10.1024/1010-0652.20.3.147.
Südkamp, Anna, Johanna Kaiser, und Jens Möller. 2012. «Accuracy of Teachers’ Judgments of Students’ Academic Achievement: A Meta-Analysis». Journal of Educational Psychology 104 (März):743–62. https://doi.org/10.1037/a0027627.
Swiecki, Zachari, Hassan Khosravi, Guanliang Chen, Roberto Martinez-Maldonado, Jason M. Lodge, Sandra Milligan, Neil Selwyn, und Dragan Gašević. 2022. «Assessment in the age of artificial intelligence». Computers and Education: Artificial Intelligence 3 (Januar):100075. https://doi.org/10.1016/j.caeai.2022.100075.
Tharwat, Alaa. 2018. «Classification Assessment Methods: a detailed tutorial», August. https://doi.org/10.1016/j.aci.2018.08.003.Wilkens, Robert. 2020. «Bewerten ohne Klausur: Kompetenzorientierte, semesterbegleitende Leistungsmessung Studierender.» die hochschullehre 6: 499–503. https://doi.org/10.3278/HSL2038W.
Zhang, Ke, und Ayse Begum Aslan. 2021. «AI technologies for education: Recent research & future directions». Computers and Education: Artificial Intelligence 2 (Januar):100025. https://doi.org/10.1016/j.caeai.2021.100025.
