目的 利用人工智能深度学习技术预测医师资格考试试题的难度,准确控制试卷难度。方法 利用构建属性模型与语义模型进行试题难度的预估,并将预估结果和专家预估结果与实测难度分别进行相关分析和重复测量方差分析,以评价采用模型进行医学试题难度预估的可行性和有效性。结果 对于某年整卷试题难度预估,属性模型预估结果与实测难度的皮尔森相关系数为0.266,略低于专家预估难度与实测难度的相关系数0.356,2个系数置信区间有交叉,差异无统计学意义(P>0.05);语义模型预估结果与实测难度的皮尔森相关系数为0.512,高于专家预估难度与实测难度的相关系数0.356,2个系数置信区间无交叉,差异具有统计学意义(P<0.05)。重复测量方差分析发现,仅语义模型预估难度与实测难度的差异无统计学意义(P>0.05)。结论 使用语义模型预估的试题难度比专家预估的难度更接近实测难度,可以尝试将该方法在考前应用于试题难度预估,结合专家预估的结果共同指导组卷,从而更加客观、准确地把握试卷难度。
Objective To predict the difficulty of questions used by physician qualifying examination and to accurately control the difficulty of the whole paper by using artificial intelligence (AI) deep learning technology. Methods By constructing the framework of attribute model and semantic model, the difficulty of the test questions was estimated. The results of AI prediction and experts' prediction were correlated and repeated-measures ANOVA were conducted with the actual test difficulty respectively for evaluation of feasibility and effectiveness of the AI model applied for difficulty prediction. Results For a given year's whole paper questions, the Pearson correlation coefficient between attribute model prediction and actual test difficulty was 0.266, which was slightly lower than the correlation coefficient of 0.356 between the experts prediction difficulty and the actual difficulty. There was a crossover between the confidence intervals of the two coefficients (P>0.05). The Pearson correlation coefficient between semantic model and actual difficulty was 0.512, which was higher than the correlation coefficient between the difficulty by experts' prediction and the actual test (0.356). There was no crossover in the confidence intervals of the two coefficients (P<0.05). The results of statistical analysis using one-way repeated measures ANOVA showed that there was no statistical difference only between sematic model prediction difficulty and actual test difficulty (P>0.05). Conclusions The difficulty of the test questions predicted by the semantic model is closer to the actual difficulty of the test than the difficulty predicted by the experts. So it may be applied to the pre-examination difficulty prediction, and combined with the results of the experts' prediction to jointly guide the development of test paper based a well-predicted difficulty.
[1] 曹开奉,王伟群,刘芳.我国高考理科试题难度影响因素的文献分析[J].考试研究,2018(3):38-44, 31.
[2] 杨涛,辛涛,杨婷婷.试题难度的主观预估方法[J].中国考试,2014(2):3-9.
[3] Hsu FY, Lee HM, Chang TH, et al. Automated estimation of item difficulty for multiple-choice tests: an application of word embedding techniques[J]. Information Processing and Management, 2018, 54(6): 969-984.
[4] 杨芳丽, 何佳. 医师资格考试医学综合笔试试题难度影响因素的分析[J].中华医学教育杂志,2017,37(2):312-316. DOI: 10.3760/j.issn.1673-677X.2017.02.034.
[5] 国家医学考试中心. Angoff法在中国医师资格考试医学综合笔试合格分数线确定中的应用[J].中华医学教育探索杂志,2011,10(1):87-89.
[6] 郭元祥.深度学习:本质与理念[J]. 新教师, 2017(7):11-14.
[7] 孙恒,李金波.高考试题难度的预估研究[J]. 教育理论与实践, 2008, 28(10):3-5.
[8] Huang H, Hu X, Zhao Y, et al. Modeling task fMRI data via deep convolutional autoencoder[J]. IEEE Trans Med Imaging, 2018,37(7):1551-1561. DOI: 10.1109/TMI.2017.2715285.
[9] 佟威,汪飞,刘淇,等.数据驱动的数学试题难度预测[J].计算机研究与发展,2019,56(5):1007-1019.
[10] 李洪福,李振来.试题难度预估方法的探索与实践[J].生物学通报,2006,41(7):43-45. DOI: 10.3969/j.issn.0006-3193.2006.07.024.
[11] 陈灵芝,丁晓娟,余莉,等.医学微生物学试题难度预估方法的探索[J].医学教育探索,2009, 8(8):996-998.
[12] LeCun Y, Bottou L, BENGio Y, et al. Gradient-based learning applied to document recognition[J]. Proc of the IEEE, 1998, 86(11):2278-2324. DOI: 10.1109/5.726791.
[13] Wu J, Liu X, Zhang X, et al. Master clinical medical knowledge at certificated-doctor-level with deep learning model[J]. Nat Commun, 2018,9(1):4352. DOI: 10.1038/s41467-018-06799-6.
[14] Kotsiantis SB. Supervised machine learning: a review of classification techniques[J]. Informatica, 2007 (31) : 249-268.