目的 评价临床医学专业(本科)水平测试技能考试信度,为改进技能考试质量提供依据。方法 应用概化理论对2020年完成临床医学专业(本科)水平测试技能考试的7 322名考生的成绩进行研究。根据考试设计,纳入的因素包括考生、考站和考题,计算各因素的方差百分比,分析影响考试信度的主要因素,计算考试的概括力系数和可靠性指数分析考试的信度。基于完成同一套试题的考生分数,通过调整考题和考站数量,得出不同概括力系数和可靠性指数以探讨考核方案的优化。结果 考题的方差百分比为70.0%~76.0%,考站的方差百分比为17.6%~24.4%,考生的方差百分比为0.3%~0.8%。概括力系数范围为0.502~0.717,可靠性指数范围为0.052~0.223。当考站数量不变,考题数量增加1道时,概括力系数增加0.002,可靠性指数增加0.002,考题数量增加5道时,概括力系数增加0.007,可靠性指数增加0.004;当考站数量增加1站时,概括力系数增加0.025,可靠性指数增加0.010,考站数量增加2站时,概括力系数增加0.045,可靠性指数增加0.017。结论 水平测试临床基本技能考试质量尚存在提升空间,考站与考题是影响测试信度的主要因素,通过增加考站与考题数量可以提高测试信度,但需综合考虑测试时间及考站、考题开发成本。
Objective To assess the reliability of the clinical fundamental skills test forstandardized competence test for clinical medicine undergraduates and provide an empirical basis for improving the quality of the test. Methods A generalizability theory was employed to examine the scores of 7 322 examinees who completed the clinical fundamental skills test for BSc in 2020. According to the test design, examinees, stations and items were included in the analysis.The percentage of variance for each facet was calculated to identify the main sources affecting test reliability. Generalizability (G) coefficient and phi index were calculated to analyze the reliability of the test. Based on the scores of examinees who completed the same items in the test, different G coefficients and phi indices were obtained to explore the optimization of the test reliability by adjusting the number of items and stations. Results The results from the generalizability theory showed that the percentage of variance for items was 70.0%~76.0%, the percentage of variance for stations was 17.6%~24.4%, and the percentage of variance for examinees was 0.3%~0.8%. G coefficients for the test ranged from 0.502 to 0.717, and phi indices ranged from 0.052 to 0.223. When one item was added, G coefficient increased by 0.002 and phi index increased by 0.002. When five items were added, G coefficient increased by 0.007 and phi index increased by 0.004. When one station was added, G coefficient increased by 0.025, and phi index increased by 0.010. When two stations were added, G coefficient increased by 0.045 and phi index increased by 0.017. Conclusions There is still room for improving the quality of clinical fundamental skills test. The stations and items are the main sources affecting the test reliability. The test reliability can be improved by increasing the number of stations and items, but the time and the cost of stations and items need to be considered in a comprehensive manner.
[1] 临床医学专业(本科)水平测试简介.国家医学考试网.[EB/OL]. (2021-06-03) [2021-06-03].https://www.nmec.org.cn/Pages/ArticleInfo-13-11592.html.
[2] Halman S, Fu A, Pugh D. Entrustment within an objective structured clinical examination (OSCE) progress test: bridging the gap towards competency-based medical education[J]. Med Teach, 2020,42(11):1283-1288. DOI: 10.1080/0142159X.2020.1803251.
[3] Oliveira JM, Spositon T, Cerci Neto A, et al. Functional tests for adults with asthma: validity, reliability, minimal detectable change, and feasibility[J]. J Asthma, 2022,59(1):169-177. DOI: 10.1080/02770903.2020.1838540.
[4] Bilgic E, Watanabe Y, McKendy KM, et al. Reliable assessment of performance in surgery: a practical approach to generalizability theory[J]. J Surg Educ, 2015,72(5):774-775. DOI: 10.1016/j.jsurg.2015.04.020.
[5] Brennan RL. Generalizability theory[M]. New York: Springer-Verlag, 2001:53-140.
[6] Andersen S, Nayahangan LJ, Park YS, et al. Use of generalizability theory for exploring reliability of and sources of variance in assessment of technical skills: a systematic review and meta-analysis[J]. Acad Med, 2021,96(11):1609-1619. DOI: 10.1097/ACM.0000000000004150.
[7] Clauser BE, Balog K, Harik P, et al. A multivariate generalizability analysis of history-taking and physical examination scores from the USMLE step 2 clinical skills examination[J]. Acad Med, 2009,84(10 Suppl):86-89. DOI: 10.1097/ACM.0b013e3181b36fda.
[8] 于惊涛. 我国临床执业医师资格考试的信度和效度研究[J].中国卫生事业管理,2008,25(5):349-350,354. DOI: 10.3969/j.issn.1004-4663.2008.05.029.
[9] 高超, 张革来, 王震, 等. 临床执业医师考试质量评价——以浙江考区为例[J].浙江医学教育,2021,20(2):4-6,22. DOI: 10.3969/j.issn.1672-0024.2021.02.002.
[10] Monteiro S, Sullivan GM, Chan TM. Generalizability theory made simple(r): an introductory primer to G-studies[J]. J Grad Med Educ, 2019,11(4):365-370. DOI: 10.4300/JGME-D-19-00464.1.
[11] Li G, Xie J, An L, et al. A generalizability analysis of the mobile phone addiction tendency scale for Chinese college students[J]. Front Psychiatry, 2019,10:241. DOI: 10.3389/fpsyt.2019.00241.
[12] Peeters MJ, Cor MK. Guidance for high-stakes testing within pharmacy educational assessment[J]. Curr Pharm Teach Learn, 2020,12(1):1-4. DOI: 10.1016/j.cptl.2019.10.001.
[13] 邹扬, 缪青, 芦开芳, 等. 本科和长学制毕业考试中客观结构化临床考试的应用[J].上海交通大学学报(医学版),2008, 28(z1): 71-75.
[14] Corrie K, Wiles M, Flack J, et al. A novel objective structured clinical examination for the assessment of transfusion practice in anaesthesia[J]. Clin Teach, 2011,8(2):97-100. DOI: 10.1111/j.1743-498X.2010.00404.x.
[15] Brannick MT, Erol-Korkmaz HT, Prewett M. A systematic review of the reliability of objective structured clinical examination scores[J]. Med Educ, 2011,45(12):1181-1189. DOI: 10.1111/j.1365-2923.2011.04075.x.
[16] Norman G, Bordage G, Page G, et al. How specific is case specificity?[J]. Med Educ, 2006,40(7):618-623. DOI: 10.1111/j.1365-2929.2006.02511.x.
[17] Clauser BE, Harik P, Margolis MJ, et al. The generalizability of documentation scores from the USMLE Step 2 Clinical Skills examination[J]. Acad Med, 2008,83(10 Suppl):S41-44. DOI: 10.1097/ACM.0b013e318183cd1d.
[18] Dory V, Gagnon R, Charlin B. Is case-specificity content-specificity? an analysis of data from extended-matching questions[J]. Adv Health Sci Educ Theory Pract, 2010,15(1):55-63. DOI: 10.1007/s10459-009-9169-z.
[19] Tsai TH, Shin CD, Neumann LM, et al. Generalizability analyses of NBDE Part II[J]. Eval Health Prof, 2012,35(2):169-181. DOI: 10.1177/0163278711425382.