PubMed دسترسی آزاد

Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Medical education faces increasing demand for scalable grading solutions. Large language models have been proposed as automated graders for high-stakes assessments, but evidence for their reliability in multimodal practical examinations remains limited. METHODS: This retrospective inter-rater reliability study compared three LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku with original human reference scores on 19 integrated anatomy, histology, and physiology objective structured practical examination items completed by 309 pre-medical students. Human reference scores were assigned during live summative grading by two content experts who divided items, with each item marked by one expert across all students; no human inter-rater reliability estimate was available. Items required simultaneous image interpretation and short-answer responses. Models were evaluated using zero-shot prompting reflecting deployment-realistic conditions. Agreement was assessed using Spearman's ρ, Cohen's κ, intraclass correlation coefficients, and Bland-Altman analysis. Item-level gap analysis compared student success rates with LLM performance across items. RESULTS: Rank-order correlations were strong across all models (ρ = 0.784-0.921), but categorical agreement diverged substantially. Pass/fail agreement ranged from 52.1% (ChatGPT; κ = 0.067, slight) to 84.1% (Claude; κ = 0.680, substantial). Bland-Altman analysis showed inconsistent systematic bias: Claude over-scored by +2.5 points, Gemini under-scored by -5.0 points, and ChatGPT by -10.0 points. Item-level gap analysis highlighted three recurring divergence patterns: non-standard histological staining, three-dimensional spatial reasoning from two-dimensional images, and semantic inflexibility in short-answer evaluation. The most extreme case was a pelvic three-dimensional model item on which 93.3% of students succeeded by human grading but all three LLMs assigned a mean score of zero. CONCLUSIONS: Strong rank-order correlation alone does not support LLM grading for categorical decisions in high-stakes assessment contexts. LLMs may serve as assessment assistants for first-pass ranking and discrepancy flagging, but item-level human review of flagged discordances, error-profile monitoring, and preserved human authority over pass/fail decisions are required before high-stakes deployment.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

Large language modelsOSPEautomated gradinginter- responsible AImultimodal assessment
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

[THE SECONDARY MEDICAL EDUCATION IN RUSSIA: CHALLENGES OF PRACTICAL TRAINING AND STRATEGIES ENHANCING COMPETITIVE ABILITIES OF GRADUATES AT LABOR MARKET].

The article analyzes current state of system of secondary vocational medical education in Russia. On the basis of data from government agencies and professional associations key issues are considered: record-breaking outflow of young professionals from industry, irregularity of practical training and low efficiency of existing mechanisms of employment. The particular attention is paid to strategies of increasing competitiveness of grad…

PubMed2026

Improving Resident Knowledge of Artificial Intelligence Ethics and Prompting for Clinical Use.

BACKGROUND: The rapid introduction of AI into clinical practice shifts how we must teach resident trainees so they may become ethical patient-facing clinicians in an AI-integrated healthcare system. Currently, few published innovations assess outcomes beyond learner attitudes. We developed a pilot curricular innovation to equip postgraduate Internal Medicine resident trainees with the attitudes and knowledge needed to responsibly integ…

PubMed2026

Integrated Team-Based Learning in a UK Undergraduate Medical Programme.

BACKGROUND: There is limited published evidence supporting integrated team-based learning (TBL) as an effective method for teaching undergraduate medical students. This study describes student and staff perceptions, assessment outcomes and financial factors after integrated TBL was implemented into Year 1 of a large UK undergraduate medical programme. METHODS: Five methods of data collection were used. An online survey was distributed …

PubMed2026

Integrating Spiritual Care Teaching Within the Australian Medical Curriculum.

BACKGROUND: Spiritual care has been shown to be an important component of holistic patient care. However, students have reported it missing from current Australian medical school curricula. The aim of this study was to evaluate the impact of a three-hour spiritual care workshop designed to enable final year medical students to take a spiritual history from their patients. METHODS: We used a prospective pilot study design to evaluate a …