Chinese Journal of Medical Education ›› 2026, Vol. 46 ›› Issue (8): 615-621.DOI: 10.3760/cma.j.cn115259-20250815-00959

• Medical Education Assessment • Previous Articles     Next Articles

Construction and empirical research on a BOPPPS and large language models (LLMs) integrated teaching quality evaluation system

Ou Jian1, Hu Yueqiang1, Huang Hongna1, Luo Jun2, Wang Chunling3, Liang Jiyao1   

  1. 1Teaching Department, The First Affiliated Hospital of Guangxi University of Chinese Medicine, Nanning 530015, China;
    2General Section, Information Technology Center, Guangxi University of Chinese Medicine, Nanning 530200, China;
    3Educational Teaching Quality Supervision Department, Educational Teaching Quality Supervision and Evaluation Center, Guangxi University of Chinese Medicine, Nanning 530200, China
  • Received:2025-08-15 Published:2026-07-29
  • Contact: Huang Hongna, Email: 306211198@qq.com
  • Supported by:
    2025 Guangxi Higher Education Undergraduate Teaching Reform Project (2025JGZ147); 2025 Guangxi Higher Education Undergraduate Teaching Reform Project (2025JGZ149); 2020 Teaching Reform Project of the First Clinical Medical College of Guangxi University of Chinese Medicine (2020JG02B)

Abstract: Objective To construct a BOPPPS-Large Language Models (LLMs) integrated teaching quality evaluation system, verify its scientific rigor and applicability, and inform the establishment of similar digital teaching quality evaluation systems in the field of traditional Chinese medicine (TCM). Methods With the BOPPPS teaching model as the structural anchor and LLMs as the quantitative support, an evaluation system covering five core dimensions of core cognition, teaching process, teaching effect, non-cognitive ability, and long-term migration was constructed. A cluster randomized controlled trial was conducted from September 2023 to June 2024, and 120 undergraduate students of the 2021 cohort majoring in TCM who took the Traditional Chinese Internal Medicine course at Guangxi University of Chinese Medicine enrolled in the study. They were equally divided into the experimental group and the control group by random number table method, 60 students in the experimental group adopted the BOPPPS-LLMs integrated evaluation system, and 60 students in the control group continued with the traditional experience-based evaluation system. Independent samples t-test was used for inter-group comparison, Pearson correlation analysis was adopted to analyze the correlation between indicators, and Cronbach′s α coefficient was applied to test the internal consistency reliability of the system. Results The overall Cronbach′s α coefficient of the evaluation system was 0.86, indicating good internal consistency reliability. All indicators of the five core dimensions in the experimental group were significantly better than those in the control group (all P<0.001), for example, the confusion degree of case/text interpretation in the core cognition dimension was (18.5±3.8) points, which was significantly lower than (32.1±4.5) points in the control group; the average score of the final comprehensive assessment of the teaching effect dimension in the experimental group was (89.5±4.2) points, which was significantly higher than (77.2±5.1) points in the control group. In addition, there was a strong negative correlation between core cognitive confusion score and teaching effect (|r|≥0.78, P < 0.001), which confirmed the good construct validity of the system. Conclusions The BOPPPS-LLMs integrated teaching quality evaluation system is scientific, reliable and highly practical. It can accurately and comprehensively evaluate the teaching quality of TCM courses, and provide an empirical basis for the implementation of digital teaching quality evaluation systems in TCM higher education.

Key words: Artificial intelligence, Educational model, Educational evaluation, Traditional Chinese medicine

CLC Number: