Development and Content Validity Evidence of an Instrument to Assess Students’ Perceptions of an AI-Enhanced Virtual Patient.
Abstract
The incorporation of artificial intelligence-enhanced virtual patients (AI-VP) in medical education requires instruments that rigorously assess students’ perceptions of their usefulness, ease of use, realism and educational value. A questionnaire was developed to assess students’ experience with a virtual patient named María V., designed based on OSCE materials. To analyze the initial evidence of content validity of an instrument designed to assess students’ perceptions of an artificial intelligence-enhanced virtual patient. Methods: A methodological content validity study was conducted. The initial instrument consisted of 23 items distributed across theoretically defined dimensions linked to the Technology Acceptance Model and simulation-based medical education: educational value, quality of clinical interaction, usability and technical functioning, realism and characterization, overall satisfaction and open-ended questions. Thirteen expert judges evaluated the relevance of each item using a four-point Likert-type scale. The item-level content validity index, scale-level content validity index based on the average method, universal agreement index and modified kappa were calculated. Results: Item-level content validity index values ranged from 0.69 to 1.00. Most items reached values equal to or above the recommended cut-off point. The scale-level content validity index based on the average method was 0.916, and the universal agreement index was 0.348. Modified kappa values ranged from 0.66 to 1.00. Based on quantitative results and the judges’ qualitative comments, items 6, 9 and 14 were removed, and adjustments were made to items 1, 10, 16 and 18. The questionnaire on students’ perceptions of the virtual patient María V. provides favorable initial evidence of content validity. The refinement of the instrument resulted in a final 20-item version, supported not only by statistical criteria but also by conceptual and functional coherence with the current characteristics of the virtual patient.
Downloads
-
Abstract27
-
pdf (Español (España))19
-
pdf19
-
xml (Español (España))0
References
1. Barrows HS. An overview of the uses of standardized patients for teaching and evaluating clinical skills. Acad Med. 1993, 68, 6. 443-451. https://doi.org/10.1097/00001888-199306000-00002
2. Cross J, Kayalackakom T, Robinson RE, Vaughans A, Sebastian R, Hood R, et al. Assessing ChatGPT’s capability as a new age standardized patient: qualitative study. JMIR Med Educ. 2025, 11. https://doi.org/10.2196/63353
3. Zhang B, Liu X, Wang Y, Zhou L, Xie Q, Wang B. Human or LLM as standardized patients? A comparative study for medical education [Preprint]. arXiv. 2025. https://doi.org/10.48550/arXiv.2511.14783.
4. Hamilton A, Molzahn A, McLemore K. The evolution from standardized to virtual patients in medical education. Cureus. 2024, 16, 10. https://doi.org/10.7759/cureus.71224
5. Messick S. Validity. En: Linn RL, editor. Educational measurement. 3rd ed. New York: Macmillan; 1989. p. 13-103.
6. Kane MT. Validation. In: Brennan RL, editor. Educational measurement. 4th ed. Westport (CT): American Council on Education/Praeger; 2006. p. 17-64.
7. Kane MT. Validating the interpretations and uses of test scores. J Educ Meas. 2013, 50, 1. 1-73. https://doi.org/10.1111/jedm.12000
8. Carrillo-Avalos BA, Leenen I, Trejo-Mejía JA, Sánchez-Mendiola M. Bridging validity frameworks in assessment: beyond traditional approaches in health professions education. Teach Learn Med. 2025, 37, 2. 229-238. https://doi.org/10.1080/10401334.2023.2293871
9. Downing SM. Validity: on the meaningful interpretation of assessment data. Med Educ. 2003, 37, 9. 830-837. https://doi.org/10.1046/j.1365-2923.2003.01594.x
10. Van der Vleuten CPM, Schuwirth LWT, Driessen EW, Dijkstra J, Tigelaar D, Baartman LKJ, et al. A model for programmatic assessment fit for purpose. Med Teach. 2012, 34, 3. 205-214. https://doi.org/10.3109/0142159X.2012.652239
11. Briz-Ponce L, García-Peñalvo FJ. An empirical assessment of a technology acceptance model for apps in medical education. J Med Syst. 2015, 39, 176. https://doi.org/10.1007/s10916-015-0352-x
12. Lee JWY, Ong DW, Soh RCC, Rao JP, Bello F. Technology Acceptance Model in medical education: systematic review. JMIR Med Educ. 2025, 11. https://doi.org/10.2196/67873
13. Lynn MR. Determination and quantification of content validity. Nurs Res. 1986, 35, 6. 382-385.
14. Davis FD. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Q. 1989, 13, 3. 319-340. https://doi.org/10.2307/249008
15. Cook DA, Erwin PJ, Triola MM. Computerized virtual patients in health professions education: a systematic review and meta-analysis. Acad Med. 2010, 85, 10. 1589-1602. https://doi.org/10.1097/ACM.0b013e3181edfe13
16. INACSL Standards Committee. Healthcare Simulation Standards of Best Practice®. Clin Simul Nurs. 2021, 58. 1-93. https://doi.org/10.1016/j.ecns.2021.08.018
17. Zeng J, Qi W, Shen S, Liu X, Li S, Wang B, et al. Embracing the future of medical education with large language model-based virtual patients: scoping review. J Med Internet Res. 2025, 27. https://doi.org/10.2196/79091.
18. Polit DF, Beck CT, Owen SV. Is the CVI an acceptable indicator of content validity? Appraisal and recommendations. Res Nurs Health. 2007, 30, 4. 459-467. https://doi.org/10.1002/nur.20199
19. Wynd CA, Schmidt B, Schaefer MA. Two quantitative approaches for estimating content validity. West J Nurs Res. 2003, 25, 5. 508-518. https://doi.org/10.1177/0193945903252998
20. Fleiss JL. Measuring nominal scale agreement among many raters. Psychol Bull. 1971, 76, 5. 378-382. https://doi.org/10.1037/h0031619
21. Cicchetti DV, Sparrow SA. Developing criteria for establishing interrater reliability of specific items: application to assessment of adaptive behavior. Am J Ment Defic. 1981, 86, 2. 127-137.
22. Grevisse C, et al. ECOSBot: a generative AI-based tool designed to simulate both standardized patients and standardized examiners in nephrology-focused OSCEs. Clin Kidney J. 2025, 18, 10. https://doi.org/10.1093/ckj/sfaf308
23. Borg A, Schiött J, Ivegren W, Gentline C, Huss V, Hugelius AM, et al. AI-generated feedback following social robotic virtual patient interactions and medical student performance: nonrandomized quasi-experimental study. JMIR Med Educ. 2026, 12. https://doi.org/10.2196/90368
24. Gamble C, Oatham A, Parikh R. Should virtual OSCE teaching replace or complement face-to-face teaching in the post-COVID-19 educational environment. Cureus. 2023, 15, 11. https://doi.org/10.7759/cureus.49708
25. Khojasteh L, Karimian Z, Nasiri E, Kafipour R, Farahmandi AY. Artificial intelligence and academic writing questionnaire (AI-AWQ): development and validation among medical students' experiences using exploratory factor analysis. BMC Med Educ. 2025, 25, 1. 1697. https://doi.org/10.1186/s12909-025-08288-z
26. Yi Y, Kim KJ. The feasibility of using generative artificial intelligence for history taking in virtual patients. BMC Res Notes. 2025, 18, 1. 80. https://doi.org/10.1186/s13104-025-07157-8
27. American Educational Research Association, American Psychological Association, National Council on Measurement in Education. Standards for educational and psychological testing. Washington (DC): American Educational Research Association; 2014.
Copyright (c) 2026 Servicio de Publicaciones de la Universidad de Murcia

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
The works published in this magazine are subject to the following terms:
1. The Publications Service of the University of Murcia (the publisher) preserves the economic rights (copyright) of the published works and favors and allows them to be reused under the use license indicated in point 2.
2. The works are published under a Creative Commons Attribution-NonCommercial-NoDerivative 4.0 license.
3. Self-archiving conditions. Authors are allowed and encouraged to disseminate electronically the pre-print versions (version before being evaluated and sent to the journal) and / or post-print (version evaluated and accepted for publication) of their works before publication , since it favors its circulation and earlier diffusion and with it a possible increase in its citation and reach among the academic community.