中华医学教育杂志 ›› 2026, Vol. 46 ›› Issue (8): 628-635.DOI: 10.3760/cma.j.cn115259-20250729-00842

• 医学教育评估 • 上一篇    下一篇

生成式大语言模型在提升基础医学试题质量中的应用效果

刘侃1, 郑淑园1, 王振禹2, 刘子冬2, 杨海2, 田甜3   

  1. 1空军军医大学基础医学院外语教研室,西安 710032;
    2空军军医大学基础医学院教学科研处,西安 710032;
    3空军军医大学口腔医院特诊科,西安 710032
  • 收稿日期:2025-07-29 发布日期:2026-07-29
  • 通讯作者: 田甜, Email: 50472849@qq.com
  • 基金资助:
    2025年空军军医大学基础医学院教学研究课题(ZL2025007)

Application of generative large language models in improving the quality of basic medical science examination items

Liu Kan1, Zheng Shuyuan1, Wang Zhenyu2, Liu Zidong2, Yang Hai2, Tian Tian3   

  1. 1Department of Foreign Languages, Basic Medical Science Academy, Air Force Medical University, Xi′an 710032, China;
    2Office of Teaching Affairs and Scientific Research, Basic Medical Science Academy, Air Force Medical University, Xi′an 710032, China;
    3Special Diagnosis Department of the Third Affiliated Hospital, Air Force Medical University, Xi′an 710032, China
  • Received:2025-07-29 Published:2026-07-29
  • Contact: Tian Tian, Email: 50472849@qq.com
  • Supported by:
    Teaching Research Project of the Basic Medical Science Academy, Air Force Medical University, 2025 (ZL20250007)

摘要: 目的 探讨生成式大语言模型(large language model,LLM)辅助提升基础医学试题质量的效果。方法 选取空军军医大学2013至2024年“基础医学综合测试”试题2 520道,依据答对率、区分度和干扰项选择率从中筛选出64道存在问题的试题。采用思维树框架设计3套提示词工程,通过ChatGPT-4o模型生成121道新试题,经过专家组人工调优后随机抽取48题编入实际考试进行实测。采用Wilcoxon符号秩检验比较答对率和区分度的差异,采用Fisher精确检验比较干扰项选择率的差异。结果 生成试题的平均答对率由0.218 1提升至0.444 2(P<0.001),平均区分度由0.082 1提升至0.146 9(P=0.019)。异常干扰项结构显著改善:无选择率>30%干扰项的试题由27.1%(13/48)增至54.2%(26/48),无选择率<5%干扰项的试题由41.7%(20/48)降至33.3%(16/48),其差异均具有统计学意义(均P<0.05)。72.7%(88/121)的生成试题需要人工调优,主要涉及选项设计缺陷[44.3%(39/88)]和题干表述问题[30.7%(27/88)]。结论 LLM在基础医学试题命制中具有良好的应用潜力,通过标准化提示词工程设计和人工调优流程可以有效提升试题质量,但仍然需要专业人员深度参与质量控制。

关键词: 人工智能, 基础医学, 试题命制, 生成式大语言模型

Abstract: Objective To investigate the effectiveness of generative large language models (LLMs) in assisting the improvement of basic medical science examination item quality. Methods A total of 2 520 items from the ″Comprehensive Basic Medical Science Examination″ administered at Air Force Medical University from 2013 to 2024 were selected. Based on the correct response rate, discrimination index, and distractor selection rate, 64 problematic items were identified. Three sets of prompt engineering protocols were designed using the Tree of Thoughts framework, and 121 new items were generated through the ChatGPT-4o model. After manual refinement by an expert panel, 48 items were randomly selected and incorporated into an actual examination for field testing. The Wilcoxon signed-rank test was used to compare differences in correct response rates and discrimination indices, and Fisher′s exact test was used to compare differences in distractor selection rates. Results The mean correct response rate of the generated items increased from 0.218 1 to 0.444 2 (P<0.001), and the mean discrimination index increased from 0.082 1 to 0.146 9 (P=0.019). The abnormal distractor structure was significantly improved: the proportion of items with no distractor selection rate exceeding 30% increased from 27.1% (13/48) to 54.2% (26/48), and the proportion of items with no distractor selection rate below 5% decreased from 41.7% (20/48) to 33.3% (16/48), all differences were statistically significant (all P<0.05). A total of 72.7% (88/121) of the generated items required manual refinement, primarily involving option design deficiencies [44.3% (39/88)] and stem wording issues [30.7% (27/88)]. Conclusions LLMs demonstrate promising application potential in basic medical science item development. Standardized prompt engineering design combined with manual refinement processes can effectively improve item quality, although in-depth involvement of subject matter experts in quality control remains essential.

Key words: Artificial intelligence, Basic medical science, Test item development, Generative large language models

中图分类号: