Chinese Journal of Medical Education ›› 2026, Vol. 46 ›› Issue (8): 628-635.DOI: 10.3760/cma.j.cn115259-20250729-00842

• Medical Education Assessment • Previous Articles     Next Articles

Application of generative large language models in improving the quality of basic medical science examination items

Liu Kan1, Zheng Shuyuan1, Wang Zhenyu2, Liu Zidong2, Yang Hai2, Tian Tian3   

  1. 1Department of Foreign Languages, Basic Medical Science Academy, Air Force Medical University, Xi′an 710032, China;
    2Office of Teaching Affairs and Scientific Research, Basic Medical Science Academy, Air Force Medical University, Xi′an 710032, China;
    3Special Diagnosis Department of the Third Affiliated Hospital, Air Force Medical University, Xi′an 710032, China
  • Received:2025-07-29 Published:2026-07-29
  • Contact: Tian Tian, Email: 50472849@qq.com
  • Supported by:
    Teaching Research Project of the Basic Medical Science Academy, Air Force Medical University, 2025 (ZL20250007)

Abstract: Objective To investigate the effectiveness of generative large language models (LLMs) in assisting the improvement of basic medical science examination item quality. Methods A total of 2 520 items from the ″Comprehensive Basic Medical Science Examination″ administered at Air Force Medical University from 2013 to 2024 were selected. Based on the correct response rate, discrimination index, and distractor selection rate, 64 problematic items were identified. Three sets of prompt engineering protocols were designed using the Tree of Thoughts framework, and 121 new items were generated through the ChatGPT-4o model. After manual refinement by an expert panel, 48 items were randomly selected and incorporated into an actual examination for field testing. The Wilcoxon signed-rank test was used to compare differences in correct response rates and discrimination indices, and Fisher′s exact test was used to compare differences in distractor selection rates. Results The mean correct response rate of the generated items increased from 0.218 1 to 0.444 2 (P<0.001), and the mean discrimination index increased from 0.082 1 to 0.146 9 (P=0.019). The abnormal distractor structure was significantly improved: the proportion of items with no distractor selection rate exceeding 30% increased from 27.1% (13/48) to 54.2% (26/48), and the proportion of items with no distractor selection rate below 5% decreased from 41.7% (20/48) to 33.3% (16/48), all differences were statistically significant (all P<0.05). A total of 72.7% (88/121) of the generated items required manual refinement, primarily involving option design deficiencies [44.3% (39/88)] and stem wording issues [30.7% (27/88)]. Conclusions LLMs demonstrate promising application potential in basic medical science item development. Standardized prompt engineering design combined with manual refinement processes can effectively improve item quality, although in-depth involvement of subject matter experts in quality control remains essential.

Key words: Artificial intelligence, Basic medical science, Test item development, Generative large language models

CLC Number: