Performance of generative pre-trained transformers (GPTs) in Certification Examination of the College of Family Physicians of Canada

作　　者：Mehdi Mousavi Shabnam Shafiee Jason M Harley Jackie Chi Kit Cheung Samira Abbasgholizadeh Rahimi

出　　处：《Family Medicine and Community Health》2024年第S01期12-19,共8页家庭医学与社区卫生（英文）

基　　金：SAR is Canada Research Chair(Tier Ⅱ)in Advanced Digital Primary Health Care,received salary support from a Research Scholar Junior 1 Career Development Award from the Fonds de Recherche du Québec-Santé(FRQS)during a portion of this study,and her research program is supported by the Natural Sciences Research Council(NSERC)Discovery(grant 2020-05246).

摘　　要：Introduction The application of large language models such as generative pre-trained transformers(GPTs)has been promising in medical education,and its performance has been tested for different medical exams.This study aims to assess the performance of GPTs in responding to a set of sample questions of short-answer management problems(SAMPs)from the certification exam of the College of Family Physicians of Canada(CFPC).Method Between August 8th and 25th,2023,we used GPT-3.5 and GPT-4 in five rounds to answer a sample of 77 SAMPs questions from the CFPC website.Two independent certified family physician reviewers scored AI-generated responses twice:first,according to the CFPC answer key(ie,CFPC score),and second,based on their knowledge and other references(ie,Reviews’score).An ordinal logistic generalised estimating equations(GEE)model was applied to analyse repeated measures across the five rounds.Result According to the CFPC answer key,607(73.6%)lines of answers by GPT-3.5 and 691(81%)by GPT-4 were deemed accurate.Reviewer’s scoring suggested that about 84%of the lines of answers provided by GPT-3.5 and 93%of GPT-4 were correct.The GEE analysis confirmed that over five rounds,the likelihood of achieving a higher CFPC Score Percentage for GPT-4 was 2.31 times more than GPT-3.5(OR:2.31;95%CI:1.53 to 3.47;p<0.001).Similarly,the Reviewers’Score percentage for responses provided by GPT-4 over 5 rounds were 2.23 times more likely to exceed those of GPT-3.5(OR:2.23;95%CI:1.22 to 4.06;p=0.009).Running the GPTs after a one week interval,regeneration of the prompt or using or not using the prompt did not significantly change the CFPC score percentage.Conclusion In our study,we used GPT-3.5 and GPT-4 to answer complex,open-ended sample questions of the CFPC exam and showed that more than 70%of the answers were accurate,and GPT-4 outperformed GPT-3.5 in responding to the questions.Large language models such as GPTs seem promising for assisting candidates of the CFPC exam by providing potential answers.However,their us

关键词：CANADA PROMPT education

分类号：TH77[机械工程—仪器科学与技术]

参考文献：

正在载入数据...

二级参考文献：

正在载入数据...

耦合文献：

正在载入数据...

引证文献：

正在载入数据...

二级引证文献：

正在载入数据...

同被引文献：

正在载入数据...

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

Performance of generative pre-trained transformers (GPTs) in Certification Examination of the College of Family Physicians of Canada

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

高级检索检索式检索

时间限定

期刊范围

学科限定全选

高级检索 检索式检索

时间限定

期刊范围

学科限定全选

Performance of generative pre-trained transformers (GPTs) in Certification Examination of the College of Family Physicians of Canada

我的收藏

参考文献：

二级参考文献：

耦合文献：

引证文献：

二级引证文献：

同被引文献：

相关期刊文献：

相关的主题

相关的作者对象

相关的机构对象

下载全文

用户登录

高级检索检索式检索