PubMed دسترسی آزاد

AI-Assisted Angoff Standard Setting for Multiple-Choice Examinations in Medical Education: Comparative Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Standard setting is essential for a defensible assessment in medical education. The modified Angoff method requires several expert judges, and determining the minimally competent candidate is cognitively challenging. However, empirical evidence on the role of AI in standard setting is unclear. OBJECTIVE: This study aimed to examine the role of large language models (LLMs) in the modified Angoff standard setting method for multiple-choice questions in a basic medical science examination compared with faculty judges. METHODS: This study was conducted in year 2 of the MD program in the United Arab Emirates. Ten faculty judges and 5 LLMs (GPT-5.2, Grok, DeepSeek, MedGemma, and Claude 4.5 Sonnet) determined the Angoff cutoff scores for a summative examination (120 multiple-choice questions). Standardized prompts were used for the LLMs to mimic the same information provided for faculty judges. Generalizability (G) theory analyses were performed using a fully crossed item × rater design to estimate variance components, G and Φ coefficients, decision study, and root mean squared error (RMSE) of the Angoff cutoff scores. RESULTS: Faculty-generated Angoff estimates (mean 69.13, SD 10.92) were comparable to LLM-generated estimates (mean 68.94, SD 12.39). A 2-tailed paired-sample t test revealed no statistically significant difference between the 2 groups (95% CI -2.21 to 2.91; t119=0.27; P=.79). Generalizability theory analysis demonstrated moderate reliability for the faculty panel (G coefficient=0.738; Φ coefficient=0.692). Despite comprising only 5 LLMs, the LLM panel demonstrated higher reliability (G coefficient=0.823; Φ coefficient=0.815), lower RMSE (1.28 vs 2.33), higher item-related variance (46.85% vs 18.36%), and lower rater-related variance (2.64% vs 16.40%) than the faculty panel. Pass rates were similar using LLM- and faculty-derived cutoff scores (52/73, 71.2%). LLMs differed in their minimally competent candidate conceptualization and approaches to determining item-level percentages of correct responses. Furthermore, the correlations between Angoff estimates and item-related P values were larger in LLMs than in faculty judges (r=0.552 vs 0.437). CONCLUSIONS: In this single-institution study, the evaluated LLMs generated modified Angoff estimates that were broadly comparable to those of faculty judges. Generalizability theory analyses demonstrated higher G and Φ coefficients, lower rater-related variance, and lower RMSE for the evaluated LLM outputs under standardized prompting conditions, indicating greater consistency in the generated estimates. Application of the resulting cutoff scores produced pass and fail rates similar to those derived from faculty judges. These findings suggest the future application of LLMs as decision assistance tools for modified Angoff standard setting while maintaining expert human oversight.

نتیجه فارسی

این مطالعه مقایسه‌ای نقش مدل‌های زبانی بزرگ (LLM) را در روش آنگوف اصلاح‌شده برای آزمون‌های چندگزینه‌ای پزشکی بررسی کرد. نتایج نشان داد که برآوردهای LLM با برآوردهای قاضیان دانشگاهی قابل مقایسه است. پنل‌های LLM پایایی بالاتری و خطای کمتری نسبت به پنل‌های دانشگاهی داشتند. نرخ موفقیت نهایی مشابهی مشاهده شد.

  • برآوردهای LLM با برآوردهای قاضیان دانشگاهی قابل مقایسه بودند.
  • پنل‌های LLM پایایی بالاتری (ضریب G و Φ) و خطای کمتر (RMSE) داشتند.
  • نرخ موفقیت نهایی با امتیازات آستانه LLM و دانشگاهیان مشابه بود.
  • LLM‌ها در مفهوم‌سازی کاندیدای حداقل متخصص و رویکردهای تصمیم‌گیری تفاوت داشتند.

ترجمه فارسی چکیده

مقدمه: تنظیم استاندارد برای ارزیابی‌های دفاع‌پذیر در آموزش پزشکی ضروری است. روش آنگوف اصلاح‌شده نیازمند چندین قاضی خبره است و تعیین کاندیدای حداقل متخصص چالش شناختی دارد. اما شواهد تجربی در مورد نقش هوش مصنوعی در تنظیم استاندارد مبهم است. هدف: بررسی نقش مدل‌های زبانی بزرگ (LLM) در روش آنگوف اصلاح‌شده برای سوالات چندگزینه‌ای در یک آزمون علوم پایه پزشکی در مقایسه با قاضیان دانشگاهی. روش: این مطالعه در سال دوم برنامه دکترای پزشکی در امارات متحده عربی انجام شد. ده قاضی دانشگاهی و ۵ مدل LLM (GPT-5.2، Grok، DeepSeek، MedGemma و Claude 4.5 Sonnet) امتیاز آستانه آنگوف را برای یک آزمون جمعی (۱۲۰ سوال چندگزینه‌ای) تعیین کردند. از پرامپت‌های استاندارد برای شبیه‌سازی اطلاعات ارائه شده به قاضیان استفاده شد. تحلیل‌های نظریه تعمیم‌پذیری با طراحی کاملاً متقاطع آیتم × قاضی برای برآورد مؤلفه‌های واریانس، ضریب G و Φ، مطالعه تصمیم و ریشه میانگین مربعات خطا (RMSE) انجام شد. نتایج: برآوردهای آنگوف تولید شده توسط دانشگاهیان (میانگین ۶۹.۱۳، انحراف معیار ۱۰.۹۲) با برآوردهای تولید شده توسط LLM (میانگین ۶۸.۹۴، انحراف معیار ۱۲.۳۹) قابل مقایسه بودند. آزمون t جفت‌الگوی دوطرفه تفاوت معنی‌دار آماری بین دو گروه را نشان نداد (۹۵٪ CI -۲.۲۱ تا ۲.۹۱؛ t119=۰.۲۷؛ P=.۷۹). تحلیل نظریه تعمیم‌پذیری نشان‌دهنده پایایی متوسط برای پنل دانشگاهیان (ضریب G=۰.۷۳۸؛ ضریب Φ=۰.۶۹۲) بود. با وجود اینکه فقط ۵ مدل LLM داشتند، پنل LLM پایایی بالاتری (ضریب G=۰.۸۲۳؛ ضریب Φ=۰.۸۱۵)، RMSE پایین‌تر (۱.۲۸ در برابر ۲.۳۳)، واریانس مرتبط با آیتم بالاتر (۴۶.۸۵٪ در برابر ۱۸.۳۶٪) و واریانس مرتبط با قاضی پایین‌تر (۲.۶۴٪ در برابر ۱۶.۴۰٪) را نشان داد. نرخ موفقیت با استفاده از امتیازات آستانه LLM و دانشگاهیان مشابه بود (۵۲/۷۳، ۷۱.۲٪). LLM‌ها در مفهوم‌سازی کاندیدای حداقل متخصص و رویکردهای تعیین درصد پاسخ‌های صحیح در سطح آیتم تفاوت داشتند. علاوه بر این، همبستگی بین برآوردهای آنگوف و مقادیر P مرتبط با آیتم در LLM‌ها بزرگتر از قاضیان دانشگاهی بود (r=۰.۵۵۲ در برابر ۰.۴۳۷). نتیجه‌گیری: در این مطالعه تک‌مؤسسه‌ای، مدل‌های LLM ارزیابی شده برآوردهای آنگوف اصلاح‌شده‌ای تولید کردند که به طور کلی با آن‌های قاضیان دانشگاهی قابل مقایسه بودند. تحلیل‌های نظریه تعمیم‌پذیری پایین‌تر واریانس مرتبط با قاضی و RMSE را برای خروجی‌های LLM تحت شرایط پرامپت‌سازی استاندارد نشان داد که نشان‌دهنده ثبات بیشتر در برآوردهای تولید شده است. کاربرد امتیازات آستانه نتیجه‌ای مشابه نرخ‌های قبول و رد را تولید کرد. این یافته‌ها نشان‌دهنده کاربرد آینده LLM‌ها به عنوان ابزار کمک تصمیم در تنظیم استاندارد آنگوف اصلاح‌شده با حفظ نظارت انسانی خبره است.

روش پژوهش

این مطالعه در سال دوم برنامه دکترای پزشکی در امارات متحده عربی انجام شد. ده قاضی دانشگاهی و ۵ مدل LLM (GPT-5.2، Grok، DeepSeek، MedGemma و Claude 4.5 Sonnet) امتیاز آستانه آنگوف را برای یک آزمون جمعی (۱۲۰ سوال چندگزینه‌ای) تعیین کردند. از پرامپت‌های استاندارد برای شبیه‌سازی اطلاعات استفاده شد.

محدودیت‌ها

این مطالعه در یک مؤسسه انجام شد و شامل ۵ مدل LLM بود. تفاوت‌های در مفهوم‌سازی کاندیدای حداقل متخصص توسط LLM‌ها و قاضیان وجود داشت.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
دانشجویان سال دوم برنامه دکترای پزشکی در امارات متحده عربی.
مداخله/مواجهه
استفاده از ۵ مدل زبانی بزرگ (GPT-5.2، Grok، DeepSeek، MedGemma و Claude 4.5 Sonnet) برای تعیین امتیاز آستانه آنگوف با استفاده از پرامپت‌های استاندارد.
مقایسه
ده قاضی دانشگاهی.
حجم نمونه
۱۰ قاضی دانشگاهی و ۵ مدل LLM.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

AIartificial intelligencelarge language modelsmedical educationmodified Angoff methodmultiple-choice examinationsstandard setting
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

[THE SECONDARY MEDICAL EDUCATION IN RUSSIA: CHALLENGES OF PRACTICAL TRAINING AND STRATEGIES ENHANCING COMPETITIVE ABILITIES OF GRADUATES AT LABOR MARKET].

The article analyzes current state of system of secondary vocational medical education in Russia. On the basis of data from government agencies and professional associations key issues are considered: record-breaking outflow of young professionals from industry, irregularity of practical training and low efficiency of existing mechanisms of employment. The particular attention is paid to strategies of increasing competitiveness of grad…

PubMed2026

Clinical Teaching Fellow Practice Within Complex Clinical-Educational Systems: An Activity Theory-Informed Case Study.

BACKGROUND: Clinical teaching occurs where patient care and education coexist. Within the United Kingdom's National Health Service, workforce pressures constrain teaching and supervision. Clinical Teaching Fellow (CTF) roles have expanded in response but are often viewed as discrete posts rather than practices situated within complex clinical-educational systems. This study examined how tensions within and across these systems shaped C…

PubMed2026

Near-Peer Anatomy-Anchored Teaching.

BACKGROUND: The transition from pre-clinical to clinical medicine is challenging, particularly in applying anatomical knowledge to patient care. Reductions in dedicated anatomy teaching time have compounded this. Near-peer teaching may help address this gap by reducing hierarchy and enhancing psychological safety, though few programmes have explicitly targeted the pre-clinical to clinical transition through the integration of anatomy w…

PubMed2026

Operating Theatre-Based Video Interventions to Enhance Student Preparation: A Scoping Review.

INTRODUCTION: Operating theatre experience is central to surgical education, yet medical students often feel unprepared due to the environment's complexity. Video-based resources are increasingly used in surgical training, but their role in improving preparedness for the operating theatre is unclear. This scoping review examined how video-based interventions influence medical students' self-reported preparedness, including confidence, …