Artificial intelligence-generated automated feedback in clinical simulation: a multisource integrative critical review of validity, learning, and transfer
Abstract
Objective: To critically analyze empirical evidence on artificial intelligence (AI)-generated automated feedback in clinical simulation for medical students, focusing on validity, educational outcomes, and transfer. Methods: A multisource integrative critical review, rather than a systematic review, was conducted following the methodological principles of Whittemore and Knafl and guidance for integrative reviews. Structured searches were performed in PubMed/MEDLINE, ERIC, and OpenAlex, supplemented by citation tracking and comparison with recent reviews. Primary studies underwent an explicit critical appraisal across design, comparator, outcome independence, reproducibility, and follow-up domains. Results: Seventeen empirical studies published between 2019 and 2026 were analyzed. High agreement with human raters was reported for some structured tasks, but specificity errors, overly generous scoring, run-to-run variability, and limitations in complex communication behaviors were also observed. Educational benefits were concentrated in immediate outcomes, while transfer to independent assessments was documented only in a subset. Conclusions: Automated feedback currently shows its clearest utility for formative and repeated practice. Available evidence does not establish general equivalence with human assessment or justify autonomous use in high-stakes summative decisions. Domain-specific validation, independent outcomes, and longitudinal follow-up are needed.
Downloads
-
Abstract0
-
pdf (Español (España))0
-
pdf0
-
xml (Español (España))0
References
1. Fierro Manrique YV, Arteaga Losada MS, Anduquia Manzano C, Escalante Ochoa AM, Novoa Álvarez RA, Charry Cuellar JD. La simulación de alta fidelidad clínica con actuación y la simulación virtual en metaverso como estrategias para la formación de futuros médicos. Rev Esp Edu Med. 2026, 5, 721571. https://revistas.um.es/edumed/article/view/721571. https://doi.org/10.6018/edumed.721571
2. Bonilla Mejia JJ, Cortés Fuenzalida TD, Polanco Aliaga DH, Martínez Carrillo C, Herrera Alcaíno AA. Inteligencia artificial en la formación clínica práctica de estudiantes de medicina de pregrado: una revisión de alcance de aplicaciones, resultados y brechas. Rev Esp Edu Med. 2026, 7(3), 710071. https://doi.org/10.6018/edumed.710071
3. Thind BS, Javidi D, Schwartz LM. Artificial intelligence in undergraduate medical education clinical skills curricula: a scoping review of implementations since 2022. Front Digit Health. 2026, 8, 1830254. https://doi.org/10.3389/fdgth.2026.1830254
4. Jiang J, Ye MZ, Kwok TTO, Wong JYH. GenAI-Supported Virtual Patients in Health Care Education: Systematic Review. J Med Internet Res. 2026, 28, e82756. https://doi.org/10.2196/82756
5. Loubbairi S, El Moussaoui Y, Lahlou L, Chakri I, Nassik H. The impact of artificial intelligence-driven simulation on the development of non-technical skills in medical education: a systematic review. J Educ Eval Health Prof. 2025, 22, 37. https://doi.org/10.3352/jeehp.2025.22.37
6. Hattie J, Timperley H. The Power of Feedback. Rev Educ Res. 2007, 77(1), 81-112. https://doi.org/10.3102/003465430298487
7. van de Ridder JMM, Stokking KM, McGaghie WC, ten Cate OTJ. What is feedback in clinical education? Med Educ. 2008, 42(2), 189-197. https://doi.org/10.1111/j.1365-2923.2007.02973.x
8. McGaghie WC, Issenberg SB, Cohen ER, Barsuk JH, Wayne DB. Does simulation-based medical education with deliberate practice yield better results than traditional clinical education? A meta-analytic comparative review of the evidence. Acad Med. 2011, 86(6), 706-711. https://doi.org/10.1097/ACM.0b013e318217e119
9. Young JQ, Van Merrienboer J, Durning S, Ten Cate O. Cognitive Load Theory: implications for medical education: AMEE Guide No. 86. Med Teach. 2014, 36(5), 371-384. https://doi.org/10.3109/0142159X.2014.889290
10. Fraser KL, Ayres P, Sweller J. Cognitive Load Theory for the Design of Medical Simulations. Simul Healthc. 2015, 10(5), 295-307. https://doi.org/10.1097/SIH.0000000000000097
11. Norcini J, Anderson B, Bollela V, Burch V, Costa MJ, Duvivier R, Galbraith R, Hays R, Kent A, Perrott V, Roberts T. Criteria for good assessment: consensus statement and recommendations from the Ottawa 2010 Conference. Med Teach. 2011, 33(3), 206-214. https://doi.org/10.3109/0142159X.2011.551559
12. Kane MT. Validating the Interpretations and Uses of Test Scores. J Educ Meas. 2013, 50(1), 1-73. https://doi.org/10.1111/jedm.12000
13. van der Vleuten CPM, Schuwirth LWT, Driessen EW, Dijkstra J, Tigelaar D, Baartman LKJ, van Tartwijk J. A model for programmatic assessment fit for purpose. Med Teach. 2012, 34(3), 205-214. https://doi.org/10.3109/0142159X.2012.652239
14. Bowers P, Graydon K, Ryan T, Lau JH, Tomlin D. Artificial intelligence-driven virtual patients for communication skill development in healthcare students: A scoping review. Australas J Educ Technol. 2024, 40(3), 39-57. https://doi.org/10.14742/ajet.9307
15. Gilbert V, Philip P, Llorca PM, Samalin L. Use of artificial intelligence-enhanced virtual patients in educational approaches to medical interview training: a systematic review. BMC Med Educ. 2026, 26, 588. https://doi.org/10.1186/s12909-026-08804-9
16. Fang EKM, Boon Y, Ho JY, Tey N, Low MJW, et al. Artificial intelligence generated virtual patients in health professions education: a scoping review. BMC Med Educ. 2026. https://doi.org/10.1186/s12909-026-10028-w
17. Mofidi SA, Khoshgoftar Z. Effectiveness of AI-enhanced virtual patients in clinical reasoning training of healthcare professionals and students: a systematic review. BMC Med Educ. 2026. https://doi.org/10.1186/s12909-026-10242-6
18. Brügge E, Ricchizzi S, Arenbeck M, Keller MN, Schur L, Stummer W, Holling M, Lu MH, Darici D. Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial. BMC Med Educ. 2024, 24, 1391. https://doi.org/10.1186/s12909-024-06399-7
19. Wang C, Li S, Lin N, Zhang X, Han Y, Wang X, Liu D, Tan X, Pu D, Li K, Qian G, Yin R. Application of Large Language Models in Medical Training Evaluation-Using ChatGPT as a Standardized Patient: Multimetric Assessment. J Med Internet Res. 2025, 27, e59435. https://doi.org/10.2196/59435
20. Snedegar R, Unger K, Craig JF, Kozlowski L, Carpenter D, Pyles E, Williamson J, Bodkins E, Williams D. Ambient Listening Devices as a Feedback Tool in the Simulated Learning Environment. Cureus. 2026, 18(2), e103222. https://doi.org/10.7759/cureus.103222
21. Nolan EJ, Burke HB. Large Language Model Evaluation and Feedback of Transcribed Simulated Physician-Patient Verbal Interactions. Front Artif Intell. 2026, 9, 1936749. https://doi.org/10.3389/frai.2026.1936749
22. Maicher KR, Zimmerman L, Wilcox B, et al. Using virtual standardized patients to accurately assess information gathering skills in medical students. Med Teach. 2019, 41(9), 1053-1059. https://doi.org/10.1080/0142159X.2019.1616683
23. Holderried F, Stegemann-Philipps C, Herrmann-Werner A, Festl-Wietek T, Holderried M, Eickhoff C, Mahling M. A Language Model-Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study. JMIR Med Educ. 2024, 10, e59213. https://doi.org/10.2196/59213
24. Yamamoto A, Koda M, Ogawa H, Miyoshi T, Maeda Y, Otsuka F, Ino H. Enhancing Medical Interview Skills Through AI-Simulated Patient Interactions: Nonrandomized Controlled Trial. JMIR Med Educ. 2024, 10, e58753. https://doi.org/10.2196/58753
25. Scherr R, Spina A, Dao A, Andalib S, Halaseh FF, Blair S, Wiechmann W, Rivera R. Novel Evaluation Metric and Quantified Performance of ChatGPT-4 Patient Management Simulations for Early Clinical Education: Experimental Study. JMIR Form Res. 2025, 9, e66478. https://doi.org/10.2196/66478
26. Weisman D, Sugarman A, Huang YM, Gelberg L, Ganz PA, Comulada WS. Development of a GPT-4-Powered Virtual Simulated Patient and Communication Training Platform for Medical Students to Practice Discussing Abnormal Mammogram Results With Patients: Multiphase Study. JMIR Form Res. 2025, 9, e65670. https://doi.org/10.2196/65670
27. Laverde N, Grévisse C, Jaramillo S, Manrique R. Integrating large language model-based agents into a virtual patient chatbot for clinical anamnesis training. Comput Struct Biotechnol J. 2025, 27, 2481-2491. https://doi.org/10.1016/j.csbj.2025.05.025
28. Suárez-García RX, Chavez-Castañeda Q, Orrico-Pérez R, et al. DIALOGUE: A Generative AI-Based Pre-Post Simulation Study to Enhance Diagnostic Communication in Medical Students Through Virtual Type 2 Diabetes Scenarios. Eur J Investig Health Psychol Educ. 2025, 15(8), 152. https://doi.org/10.3390/ejihpe15080152
29. Jacobs C, Johnson H, Tan N, Brownlie K, Joiner R, Thompson T. Application of AI Communication Training Tools in Medical Undergraduate Education: Mixed Methods Feasibility Study Within a Primary Care Context. JMIR Med Educ. 2025, 11, e70766. https://doi.org/10.2196/70766
30. Luo MJ, Bi S, Pang J, et al. A large language model digital patient system enhances ophthalmology history taking skills. NPJ Digit Med. 2025, 8(1), 502. https://doi.org/10.1038/s41746-025-01841-6
31. Byun J, Kim H, Lim J, Choi J, Ahn S. Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions. Korean J Med Educ. 2026, 38(1), 64-73. https://doi.org/10.3946/kjme.2025.108
32. Borg A, Schiött J, Ivegren W, et al. AI-generated Feedback Following Social Robotic Virtual Patient Interactions and Medical Student Performance: Nonrandomized Quasi-Experimental Study. JMIR Med Educ. 2026, 12, e90368. https://doi.org/10.2196/90368
33. Sindhvananda K, Ruengwattanachot K, Leksuwankun S, et al. AI-powered simulated patients with automated feedback for enhancing headache history-taking skills: a convergent mixed-methods study. BMC Med Educ. 2026, 26, 1151. https://doi.org/10.1186/s12909-026-09511-1
34. Chan BS, Marjot J, Dodds T, et al. Enhancing objective structured clinical examination performance through an artificial intelligence virtual patient: a proof-of-concept study. Proc (Bayl Univ Med Cent). 2026, 1-4. https://doi.org/10.1080/08998280.2026.2699604
35. Cook DA, Overgaard J, Pankratz VS, Del Fiol G, Aakre CA. Virtual Patients Using Large Language Models: Scalable, Contextualized Simulation of Clinician-Patient Dialogue With Feedback. J Med Internet Res. 2025, 27, e68486. https://doi.org/10.2196/68486
36. Lee J, Kim H, Kim KH, Jung D, Jowsey T, Webster CS. Effective virtual patient simulators for medical communication training: a systematic review. Med Educ. 2020, 54(9), 786-795. https://doi.org/10.1111/medu.14152
37. Kononowicz AA, Woodham LA, Edelbring S, et al. Virtual Patient Simulations in Health Professions Education: Systematic Review and Meta-Analysis by the Digital Health Education Collaboration. J Med Internet Res. 2019, 21(7), e14676. https://doi.org/10.2196/14676
38. Durning SJ, Artino AR Jr. Situativity theory: a perspective on how participants and the environment can interact: AMEE Guide No. 52. Med Teach. 2011, 33(3), 188-199. https://doi.org/10.3109/0142159X.2011.550965
39. Hippe DS, Umoren RA, McGee A, Bucher SL, Bresnahan BW. A targeted systematic review of cost analyses for implementation of simulation-based education in healthcare. SAGE Open Med. 2020, 8, 2050312120913451. https://doi.org/10.1177/2050312120913451
40. Masters K. Ethical use of Artificial Intelligence in Health Professions Education: AMEE Guide No. 158. Med Teach. 2023, 45(6), 574-584. https://doi.org/10.1080/0142159X.2023.2186203
41. Cantillon P, Sargeant J. Giving feedback in clinical settings. BMJ. 2008, 337, a1961. https://doi.org/10.1136/bmj.a1961
42. Cheng A, Eppich W, Grant V, Sherbino J, Zendejas B, Cook DA. Debriefing for technology-enhanced simulation: a systematic review and meta-analysis. Med Educ. 2014, 48(7), 657-666. https://doi.org/10.1111/medu.12432
43. Feigerlova E, Hani H, Hothersall-Davies E. A systematic review of the impact of artificial intelligence on educational outcomes in health professions education. BMC Med Educ. 2025, 25, 129. https://doi.org/10.1186/s12909-025-06719-5
44. Kıyak YS, İş-Kara T, Emekli E. Applications and Outcomes of Large-Language-Model-Generated Feedback in Undergraduate Medical Education: A Scoping Review. Med Sci Educ. 2026, 36(1), 81-99. https://doi.org/10.1007/s40670-025-02621-3
45. Çiçek FE, Ülker M, Özer M, Kıyak YS. ChatGPT versus expert feedback on clinical reasoning questions and their effect on learning: a randomized controlled trial. Postgrad Med J. 2025, 101(1199), 458-463. https://doi.org/10.1093/postmj/qgae170
46. Kıyak YS, Emekli E, İş-Kara T, Coşkun Ö, Budakoğlu Iİ. AI Teaches Surgical Diagnostic Reasoning to Medical Students: Evidence from an Experiment Using a Fully Automated, Low-Cost Feedback System. J Surg Educ. 2025, 82(10), 103639. https://doi.org/10.1016/j.jsurg.2025.103639
47. Kıyak YS, Coşkun Ö, Budakoğlu Iİ. 'ChatGPT can make mistakes' warnings fail: A randomized controlled trial. Med Educ. 2026, 60(2), 138-142. https://doi.org/10.1111/medu.70056
54. Elizondo-García J. Vibe-coding como competencia docente en el uso de Inteligencia Artificial para la Educación Médica: una revisión de alcance. Rev Esp Edu Med. 2026, 7(2). https://revistas.um.es/edumed/article/view/705121, https://doi.org/10.6018/edumed.705121
53. Kaya AB, Emekli E, Kiyak YS. Mapeo de aplicaciones y resultados de casos generados por modelos de lenguaje amplio en la educación de profesiones de la salud: una revisión de alcance. Rev Esp Edu Med. 2026, 7(1). https://revistas.um.es/edumed/article/view/691691, https://doi.org/10.6018/edumed.691691
52. Jiménez Salcedo SC, Granda Cruz CA. Uso de ChatGPT como herramienta pedagógica en la educación médica: revisión sistemática de la literatura (2020-2025). Rev Esp Edu Med. 2026, 7(4). https://doi.org/10.6018/edumed.716681
51. Canabal Berlanga A, Sánchez Giralt JA, Abad Santamaría B. Actualidad en la educación médica con la inteligencia artificial generativa. Rev Esp Edu Med. 2025, 6(6). https://doi.org/10.6018/edumed.684461
50. Hong QN, Fàbregues S, Bartlett G, et al. The Mixed Methods Appraisal Tool (MMAT) version 2018 for information professionals and researchers. Educ Inf. 2018, 34(4), 285-291. https://doi.org/10.3233/EFI-180221
49. Snyder H. Literature review as a research methodology: An overview and guidelines. J Bus Res. 2019, 104, 333-339. https://doi.org/10.1016/j.jbusres.2019.07.039
48. Whittemore R, Knafl K. The integrative review: updated methodology. J Adv Nurs. 2005, 52(5), 546-553. https://doi.org/10.1111/j.1365-2648.2005.03621.x
Copyright (c) 2026 Servicio de Publicaciones de la Universidad de Murcia

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
The works published in this magazine are subject to the following terms:
1. The Publications Service of the University of Murcia (the publisher) preserves the economic rights (copyright) of the published works and favors and allows them to be reused under the use license indicated in point 2.
2. The works are published under a Creative Commons Attribution-NonCommercial-NoDerivative 4.0 license.
3. Self-archiving conditions. Authors are allowed and encouraged to disseminate electronically the pre-print versions (version before being evaluated and sent to the journal) and / or post-print (version evaluated and accepted for publication) of their works before publication , since it favors its circulation and earlier diffusion and with it a possible increase in its citation and reach among the academic community.