医学教育评估

基于两种教育测量理论的全国医学博士英语统一考试质量分析

  • 张泉慧 ,
  • 张颖
展开
  • 100097北京,国家医学考试中心研究评价处

收稿日期: 2018-03-08

  网络出版日期: 2020-12-09

Quality analysis of national English test for doctor of medicine based on two kinds of educational measurement theory

  • Zhang Quanhui ,
  • Zhang Ying
Expand
  • Department of Research and Evaluation, National Medicine Examination Center, Being 100097, China

Received date: 2018-03-08

  Online published: 2020-12-09

摘要

目的 采用项目反应理论中的Rasch模型和经典测量理论,对2015年~2017年全国医学博士英语统一考试进行质量分析,比较两种理论估计结果的一致性。方法 采用资料分析方法,以2015年~2017年全国医学博士英语统一考试40 607份试卷为研究资料,采用项目反应理论的Rasch模型和经典测量理论分别分析考试的信度、难度和区分度,对比其分析结果的一致性。结果 Rasch模型的分析结果显示,3个年度试卷参数与模型拟合度较高,最大信息函数均>25.00,估计误差均<0.20,各个题型难度均在-0.70~0.70之间。经典测量理论的分析结果显示,3个年度试卷信度均>0.80,各个题型难度均在0.31~0.63之间,区分度均在0.12~0.29之间。Rasch模型和经典测量理论分析结果呈高度相关。结论 3个年度全国医学博士英语统一考试均能够较好地评价考生医学英语的应用能力。在参数估计和误差分析方面,Rasch模型的精度更高,在后续的考试评价中可以更多地使用。

本文引用格式

张泉慧 , 张颖 . 基于两种教育测量理论的全国医学博士英语统一考试质量分析[J]. 中华医学教育杂志, 2018 , 38(6) : 940 -943 . DOI: 10.3760/cma.j.issn.1673-677X.2018.06.031

Abstract

Objective To analyze the quality of national English test for doctor of medicine from 2015 to 2017 using the Rasch model of item response theory (IRT) and the classical test theory (CTT), then compare the consistency of the results. Methods 40 607 papers of national English test for doctor of medicine from 2015 to 2017 were selected as the research materials. The reliabilities, difficulties and discriminations of the examinations were analyzed using the Rasch model of IRT and CTT respectively. Then the consistency of the results was analyzed. Results IRT analysis results showed that the parameters fit well with the Rasch model, the maximum information functions were all>25.00,the estimated errors were all <0.20, and the difficulties were all between -0.70 and 0.70. CTT analysis results showed that the reliabilities of the three annual papers were all >0.80, the difficulties were all between 0.31 and 0.63, and the discriminations were all between 0.12 and 0.29. The results of the two theoretical analyses were highly correlated. Conclusions The three years' tests can evaluate the application abilities of candidates' medical English well. In terms of parameter estimation and error analysis, Rasch model is more accurate than CTT, and Rasch model should be used more in the subsequent examination evaluations.

参考文献

[1] 国家医学考试中心.全国医学博士外语统一考试指南[M].北京:人民卫生出版社,2013:1-9.
[2] 郭庆科,房洁.经典测验理论与项目反应理论的对比研究[J].山东师大学报(自然科学版),2000,15(3): 264-266.
Guo QK, Fang J. A comparative study on classical test theory and item response theory [J].Journal of Shandong Normal University(Natural Science),2000,15(3):264-266.
[3] 虞晓含.基于项目反应理论的中医体质量表条目特征研究[D].北京:北京中医药大学,2016.
[4] 韩丽珍,田琪,郝元涛.医学统计学试题库建设的构想[J].中国卫生统计,2012,29(2):220-223.
Han LZ, Tian Q, Hao YT.An Idea about the establishment of the item bank for medical statistics[J].Chinese Journal of Health Statistics,2012,29(2):220-223.
[5] 李丹,刘春华,董建鑫.基于项目反映理论的高校标准化试题库建设的探讨[J].数理医药学杂志,2011,24(6):754-755.
[6] 闫成海,杜文久,宋乃庆,等.高考数学中考试评价的研究——基于CTT与IRT的实证比较[J].华东师范大学学报(教育科学版),2014(3):10-18.
[7] 王华,陈景,马翠芹.基于CTT与IRT的试卷质量评价系统设计与实现[J].计算机工程与设计,2013,34(5):1826-1830.
Wang H, Chen J, Ma CQ. Design and implementation of paper quality assessment system based on CTT and IRT[J].Computer Engineering and Design,2013,34(5):1826-1830.
[8] 钟轶,季晓辉.两种教育测量理论在试卷质量控制和评价中的应用及其展望[J].南京医科大学学报(社会科学版),2013(1):66-71. DOI:10.7655/NYDXBSS20130118.
Zhong Y, Ji XH. Application and development of two kinds of education test theory on examination paper quality control and evaluation[J]. Acta Universitatis Medicinalis Nanjing(Social Sciences),2013 (1): 66-71. DOI:10.7655/NYDXBSS20130118.
[9] 晏子.心理科学领域内的客观测量——Rasch 模型之特点及发展趋势[J].心理科学进展,2010,18(8):1298-1305.
Yan Z.Objective measurement in psychological science:An overview of rasch model[J].Advances in Psychological Science,2010,18(8):1298-1305.
[10] 戴海崎,张锋,陈雪枫.心理教育测量[M].12版.广州:暨南大学出版社,2006:58-64.
[11] 王桂桃,严文法,田秀云.例析Rasch模型在化学试卷质量分析中的应用[J].化学教学,2016(11):14-19.
[12] 杨玉琴.化学学科能力及其测评研究[D].上海:华东师范大学,2012.
文章导航

/