PubMed دسترسی آزاد

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

ABSTRACT: BACKGROUND: Large language models (LLMs) are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist. ABSTRACT: OBJECTIVES: This study aimed to evaluate and compare the accuracy and response consistency of three leading LLMs-ChatGPT-5, Gemini 2.5, and Grok 4-on Allergy/Immunology United States Medical Licensing Examination (USMLE) Step 1-style questions under different prompt conditions. ABSTRACT: METHODS: Thirty-five USMLE Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy, and variability was measured using Shannon entropy. Mixed-effects models tested the effects of model, prompt condition, and question difficulty. ABSTRACT: RESULTS: Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance, with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability. ABSTRACT: CONCLUSION: On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Barry Bloom and the convergence of immunology, infectious disease, and public health.

In early 2026, the world lost Barry Bloom, a great advocate for public health, an extraordinarily accomplished immunologist, and a science advisor who helped refocus policy on controlling infectious diseases, including neglected diseases such as leprosy. Barry's career took him from a life of laboratory discovery where he was enormously influential in catalyzing the late 20th century shift from studying the immune response of simple mo…

PubMed2026

Trajectories in immunometabolism.

Immunometabolism has rapidly evolved from an emerging area within immunology into a topic that not only impregnates most aspects of immune research but also reveals itself as a defining feature of the immune response. At the European Immunometabolism Conference celebrated in June, we asked some of the speakers to share their personal journey as researchers in this field and what questions they are currently working on.

PubMed2026

Hazy mysteries and the major histocompatibility complex: a journey.

I was pleased to receive an invitation to write a historical perspective on my career. As I understand it, these perspectives aim to illustrate the often-twisted paths to discoveries, the fortuitous events that often enable them, and the pitfalls along the way. At the same time, they can be expected to illustrate the influence of the particular "tastes" and interests of a scientist in the choices that are made along the way and that th…