In practical educational assessments, different assessments targeting the same ability usually do not anchor to the same set of questions. Achieving score equating in the absence of anchor questions is currently a challenge that has not been fully overcome. This study proposes a low-cost solution for score equating without anchor questions based on large language models. The main principle is to utilize large language models to construct a linking group between assessments that need to be equated, thereby achieving equating. This study will take two sets of reading comprehension questions as examples, selecting GPT3.5, GPT4.0, and iFlytek Spark 3.5 to construct linking groups, and compare their effectiveness in equating from two aspects: prompt engineering (zero-shot, one-shot, and few-shot) and the number of samples in the linking group (500 and 1 000). The study found that GPT4.0 performs well in the one-shot and few-shot scenarios with a large linking group sample, indicating that the design scheme of using large language model agents as a linking group without anchor items is feasible.
在大规模的教育测评中,无锚题等值一直是一个难题。无锚题等值指的是针对完全不同的测评,将其测评分数转化到同一尺度上的解决方案[1-2]。现有的解决思路主要包括基于题目信息的等值和基于等值设计的等值,等值数据收集设计(Data Collection Design of Equating)[3]是它们的重点。然而,基于这两种思路的方案在实际实施当中均存在诸多缺陷。其中,基于题目信息的等值方案中虽然可选方法众多,但要么需要提前对题目进行精心设计[2,4],要么需要以一个大型的题目为依托[5-6],构建成本高,且实际等值效果并不理想;基于等值设计的方案,需要组织额外的测评,这往往消耗大量的人力和物力,而当待等值的测评数较多时,该方案将缺乏准确度[7]。因此在实际的无锚题等值场景中,亟需一种成本低且操作简单的等值方案。
SKAGGSG, LISSITZR W. IRT test equating: Relevant issues and a review of recent research[J]. Review of Educational Research, 1986, 56(4): 495-529. DOI: 10.3102/00346543056004495 .
[2]
ZHUW M. Test equating: What, why, how?[J]. Research Quarterly for Exercise and Sport, 1998, 69(1): 11-23. DOI: 10.1080/02701367.1998.10607662 .
WANGY B, YANGT, XINT. Progress of the study of test equating design methods with no common items[J]. Examinations Research, 2017, 13(3): 48-54 (Ch).
[5]
MISLEVYR J, SHEEHANK M, WINGERSKYM. How to equate tests with little or no data[J]. Journal of Educational Measurement, 1993, 30(1): 55-78. DOI: 10.1111/j.1745-3984.1993.tb00422.x .
[6]
LIAOC W, LIVINGSTONS A. Examining an alternative to score equating: A randomly equivalent forms approach[J]. ETS Research Report Series, 2008, 2008(1):1-22. DOI: 10.1002/j.2333-8504.2008.tb02100.x .
[7]
LIVINGSTONS A. Demographically adjusted groups for equating test scores[J]. ETS Research Report Series, 2014, 2014(2): 1-10. DOI: 10.1002/ets2.12030 .
[8]
MARENGOD, MICELIR, ROSATOR, et al. Placing multiple tests on a common scale using a post-test anchor design: Effects of item position and order on the stability of parameter estimates[J]. Frontiers in Applied Mathematics and Statistics, 2018, 4: 50. DOI: 10.3389/fams.2018.00050 .
[9]
OUYANGL, WUJ, XUJ, et al. Training language models to follow instructions with human feedback[EB/OL].[2024-02-03].
[10]
WANGS H, SUNY, XIANGY, et al. ERNIE 3.0 titan: Exploring larger-scale knowledge enhanced pre-training for language understanding and generation[EB/OL]. 2021: arXiv: 2112.12731.
[11]
BAIJ Z, BAIS, YANGS S, et al. Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond[EB/OL]. 2023: arXiv: 2308.12966.
BORDTS, VON LUXBURGU. ChatGPT participates in a computer science exam[EB/OL]. 2023: arXiv: 2303.09461. DOI: 10.1109/cvprw59228.2023.00378 .
[14]
NORIH, KINGN, McKINNEYS M, et al. Capabilities of GPT4.0 on medical challenge problems[EB/OL]. 2023: arXiv:2303.13375.
[15]
ROSENBAUMP R, RUBIND B. The central role of the propensity score in observational studies for causal effects[J]. Biometrika, 1983, 70(1): 41-55. DOI: 10.1093/biomet/70.1.41 .
[16]
TATSUOKAK K. Rule space: An approach for dealing with misconceptions based on item response theory[J]. Journal of Educational Measurement, 1983, 20(4): 345-354. DOI: 10.1111/j.1745-3984.1983.tb00212.x .
[17]
BIRENBAUMM, KELLYA E, TATSUOKAK K. Diagnosing knowledge states in algebra using the rule-space model[J]. Journal for Research in Mathematics Education, 1992, 24(5): 442-459. DOI: 10.5951/jresematheduc.24.5.0442 .
[18]
LIUQ, WUR Z, CHENE H, et al. Fuzzy cognitive diagnosis for modelling examinee performance[J]. ACM Transactions on Intelligent Systems and Technology, 2018, 9(4): Article No. 48. DOI: 10.1145/3168361 .
[19]
WANGF, LIUQ, CHENE H, et al. NeuralCD: A general framework for cognitive diagnosis[J]. IEEE Transactions on Knowledge and Data Engineering, 2023, 35(8): 8312-8327. DOI: 10.1109/TKDE.2022.3201037 .
[20]
罗照盛. 项目反应理论基础[M]. 北京: 北京师范大学出版社, 2012: 71-73.
[21]
LUOZ S. Item Response Theory[M]. Beijing: Beijing Normal University Press, 2012: 71-73 (Ch).
[22]
SEGBERSJ, SCHROEDERS. How many words do children know? A corpus-based estimation of children’s total vocabulary size[J]. Language Testing, 2017, 34(3): 297-320. DOI: 10.1177/0265532216641152 .
[23]
KORTEMEYERG. Could an artificial-intelligence agent pass an introductory physics course [J]. Physical Review Physics Education Research, 2023, 19: 010132. DOI: 10.1103/physrevphyseducres.19.010132 .
[24]
TALANT, KALINKARAY. The role of artificial intelligence in higher education: ChatGPT assessment for anatomy course[J]. Uluslararası Yönetim Bilişim Sistemleri Ve Bilgisayar Bilimleri Dergisi, 2023, 7(1): 33-40. DOI: 10.33461/uybisbbd.1244777 .
[25]
KATZD M, BOMMARITOM J, GAOS, et al. GPT-4 passes the bar exam[J]. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 2024, 382(2270): 20230254. DOI: 10.1098/rsta.2023.0254 .
WANGL. Widening the paths of assessment development via artificial intelligence-generated content: Potential applications of automatic item generation and technology-enhanced items in China[J]. Journal of China Examinations, 2023(8): 19-27. DOI: 10.19360/j.cnki.11-3303/g4.2023.08.003(Ch ).
[28]
HAEBARAT. Equating logistic ability scales by a weighted least squares method[J]. Japanese Psychological Research, 1980, 22(3): 144-149. DOI: 10.4992/psycholres1954.22.144 .
[29]
CHENJ Y, GENGY X, CHENZ, et al. Zero-shot and few-shot learning with knowledge graphs: A comprehensive survey[J]. Proceedings of the IEEE, 2023, 111(6): 653-685. DOI: 10.1109/JPROC.2023.3279374 .
ZHANGJ, RENJ. A review of evaluation criteria of equating results for the common item nonequivalent groups design[J]. China Examinations, 2018(3): 32-37. DOI: 10.19360/j.cnki.11-3303/g4.2018.03.008 (Ch ).
[32]
YAOL H. Multidimensional linking for domain scores and overall scores for nonequivalent groups[J]. Applied Psychological Measurement, 2011, 35(1): 48-66. DOI: 10.1177/0146621610373095 .