Abstract
Latent-Dirichlet-Allocation (LDA) ist eine Machine-Learning-Methode zur Themenmodellierung (engl. Topic Modelling) grosser Textkorpora (d. h. zur Bestimmung latenter Begriffs- und Dokumentencluster, sog. Themen). LDAs stellen eine transparente, ressourcenschonende und intuitiv verständliche Alternative zu LLMs dar. Im Sinne der Mixed-Methods müssen die statistischen Ergebnisse der LDAs durch qualitatives Labelling in sinnhafte Themen überführt werden. Dieses Verfahren wurde auf über 30.000 Einreichungen der European Conference on Educational Research (ECER) angewendet und liefert die Grundlage für die interaktive Webapp EduTopics (Christ et al. 2025). Sie ermöglicht Nutzenden, eigene qualitative Analyse von Themen und Ko-Variaten sowie explorative Zugänge durch interaktive Visualisierungen durchzuführen. Die Webapp und die darin enthaltenen Themen wurden in mehreren iterativen Schleifen mit Vertretenden der erziehungswissenschaftlichen Community diskutiert, angepasst und weiterentwickelt. Neben der Webapp und ihren Funktionen liegt im Beitrag ein Schwerpunkt auf zwei Granularitätsebenen mit zentralen Superthemen (k = 50) und Subthemen (k = 300). Deren hierarchische Struktur als Über- und Unterkategorien wurde über die semantische Cosinus-Ähnlichkeit der Themen bestimmt. Exemplarisch werden Trends medienpädagogisch relevanter Subthemen wie Media Literacy herangezogen und diskutiert sowie mit den Themen-Trends für das EERA-Netzwerk 06 – Open Learning: Media, Environments and Cultures (i. e. Medienpädagogik) verglichen.
Literatur
Aguaded, Igacio, Sabina Civila, und Arantxa Vizcaíno-Verdú. (2022). «Paradigm changes and new challenges for media education: Review and science mapping (2000 – 2021)». Profesional de la Información, 31(6). https://doi.org/10.3145/epi.2022.nov.06.
Aletras, Nikolaos, und Mark Stevenson. 2014. «Measuring the similarity between automatically generated topics». In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers: 22 – 27. Göteborg. https://aclanthology.org/E14-4005.pdf.
Arun, Rajkumar, Vishnu Suresh, Veni Madhavan, und Murthy Narasimha. 2010. «On finding the natural number of topics with latent Dirichlet allocation: Some observations». In Advances in Knowledge Discovery and Data Mining. PAKDD 2010. Lecture Notes in Computer Science, vol. 6118. Berlin, Heidelberg: Springer. https://doi.org/10.1007/978-3-642-13657-3_43.
Bittermann, André, und Andreas Fischer. 2018. «How to identify hot topics in psychology using topic modeling». Zeitschrift für Psychologie 226 (1): 3 – 13. https://doi.org/10.1027/2151-2604/a000318.
Blei, David. M., und John D. Lafferty. 2009. «Topic Models». In Text Mining, herausgegeben von Ashok Srivastava und Sahami Mehran, 101 – 24. New York: Chapman and Hall/CRC. https://doi.org/10.1201/9781420059458-12.
Blei, David. M., Andrew. Y. Ng, und Michael I. Jordan. 2003. «Latent Dirichlet allocation». Journal of machine Learning research 3 (Jan): 993 – 1022. https://www.jmlr.org/papers/volume3/blei03a/blei03a.pdf.
Cao, Juan, Tian Xia, Jintao Li, Yongdong Zhang, und Sheng Tang. 2009. «A density-based method for adaptive LDA model selection». Neurocomputing 72 (7 – 9): 1775 – 81. https://doi.org/10.1016/j.neucom.2008.06.011.
Christ, Alexander, Jens Röschlein, und Christoph Schindler. 2025. «EduTopics: ECER – A webapp for interactive visualisation and exploration of ECER contributions since for 1998». European Educational Research Journal. https://doi.org/10.1177/14749041251405570.
Christ, Alexander, Kathrin Smolarczyk, und Stephan Kröner. 2024. «Schwerpunktthemen der quantitativ-empirischen Forschung mit Bezug zur Digitalisierung in der kulturellen Bildung: Eine kartierende Forschungssynthese». Zeitschrift für Erziehungswissenschaft 27 (2): 351 – 92. https://doi.org/10.1007/s11618-023-01210-7.
Christ, Alexander, Marcus Penthin, und Stephan Kröner. 2021. «Big data and digital aesthetic, arts, and cultural education: Hot spots of current quantitative research». Social Science Computer Review 39 (5): 821 – 43. https://doi.org/10.1177/0894439319888455.
De Waal, Alta, und Etienne Barnard. 2008. «Evaluating topic models with stability». Cape Town: PRASA 2008. http://hdl.handle.net/10204/3016.
Deveaud, Romain, Eric SanJuan, und Patrice Bellot. 2014. «Accurate and effective latent concept modeling for ad hoc information retrieval». Document numérique 17 (1): 61 – 84. https://doi.org/10.3166/DN.17.1.61-– 84.
Fruchterman, Thomas M., und Reingold, Edward. M. 1991. «Graph drawing by force‐directed placement». Software: Practice and experience 21 (11): 1129-64. https://doi.org/10.1002/spe.4380211102.
Gandomi, Amir, und Murtaza Haider. 2015. «Beyond the hype: Big data concepts, methods, and analytics». International journal of information management 35 (2): 137 – 44. https://doi.org/10.1016/j.ijinfomgt.2014.10.007.
Griffiths, Thomas. L., und Mark Steyvers. 2004. «Finding Scientific Topics». Proceedings of the National academy of Sciences, 101(suppl_1): 5228 – 35. https://doi.org/10.1073/pnas.0307752101.
He, Qj, Bi Chen, Jian Pei, Baojun Qiu, Prasenjit Mitra, und Lee Giles. 2009. «Detecting topic evolution in scientific literature: how can citations help?». In Proceedings of the 18th ACM conference on Information and knowledge management: 957 – 66. Hong Kong: ACM. https://doi.org/10.1145/1645953.1646076.
Heidenreich, Tobias, Fabienne Lind, Jakob-Moritz Eberl, und Hajo G. Boomgaarden. 2019. «Media framing dynamics of the ‹European refugee crisis›: A comparative topic modelling approach». Journal of Refugee Studies 32 (Special_Issue_1): i172-i182. https://doi.org/10.1093/jrs/fez025.
Hobbs, Renee, und Amy Petersen Jensen. (2009). «The past, present, and future of media literacy education». Journal of media literacy education 1 (1). https://doi.org/10.23860/jmle-1-1-1.
Huang, Rui, Hailong Gai, Rong Jing, und Wan Fucheng. 2024. «Word Spectral Visualization Base on Fruchterman-Reingold Optimized ForceAtlas 2». Journal of Electrical Systems 20: 7 – 15. https://doi.org/10.52783/jes.1086.
Huang, Cui, Chao Yang, Shutao Wang, Wei Wu, Jun Su, und Chuying Liang. 2020. «Evolution of topics in education research: A systematic review using bibliometric analysis». Educational Review 72 (3): 281 – 97. https://doi.org/10.1080/00131911.2019.1566212.
Jacobs, Thomas, und Robin Tschötschel. 2019. «Topic models meet discourse analysis: a quantitative tool for a qualitative approach». International Journal of Social Research Methodology 22 (5): 469 – 85. https://doi.org/10.1080/13645579.2019.1576317.
Kobourov, Steohen G. 2012. «Spring embedders and force directed graph drawing algorithms». arXiv preprint arXiv:1201.3011. https://doi.org/10.48550/arXiv.1201.3011.
Mimno, David, Hanna M. Wallach, Edmund Talley, Miriam Leenders, und Andrew McCallum. 2011. «Optimizing semantic coherence in topic models». In Proceedings of the 2011 conference on empirical methods in natural language processing: 262 – 72. Edinburgh. https://aclanthology.org/D11-1024.Pdf.
Newman, David, Arthur Asuncion, Padhraic Smyth, und Max Welling. 2009. «Distributed algorithms for topic models». Journal of Machine Learning Research 10 (8): 1801 – 28. https://www.jmlr.org/papers/volume10/newman09a/newman09a.pdf.
O’Mara-Eves, Alison, James Thomas, John McNaught, Makoto Miwa, und Sophia Ananiadou. 2015. «Using text mining for study identification in systematic reviews: a systematic review of current approaches». Systematic reviews 4: 1 – 22. https://doi.org/10.1186/2046-4053-4-5.
Ramage, Daniel, David Hall, Ramesh Nallapati, und Christopher D. Manning. 2009. «Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora». In Proceedings of the 2009 conference on empirical methods in natural language processing: 248 – 56. https://aclanthology.org/D09-1026.pdf.
Silge, Julia, und David Robinson. 2017. Text mining with R: A tidy approach. Boston (MA): O’Reilly Media. https://www.tidytextmining.com/.
Steyvers, Mark, und Thomas L. Griffiths. 2007. «Probabilistic topic models». Handbook of Latent Semantic Analysis, 427: 424 – 40.
Stone, Matthew R. 2022. «Analysis of graph layout algorithms for use in command and control network graphs». Theses and Dissertations. 5546. https://scholar.afit.edu/etd/5546.
Wang, Xiang, Kai Zhang, Xiaoming Jin, und Dou Shen. 2009. «Mining common topics from multiple asynchronous text streams». In Proceedings of the Second ACM International Conference on Web Search and Data Mining: 192 – 201. https://doi.org/10.1145/1498759.1498826.
