PubMed دسترسی آزاد

Artificial intelligence performance in the emergency medicine subspecialty examination conducted in Türkiye.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

Emergency medicine specialists often pursue subspecialty training worldwide. In Türkiye, subspecialization in critical care medicine was introduced in March 2024, with the first entrance examination for subspecialty training in medicine (YDUS) examination having been conducted on December 15, 2024 by the Measurement, Selection, and Placement Center. Medical applications of artificial intelligence (AI), particularly GPT-4 Omni (GPT-4o), GPT-4, and Gemini-Advanced, have garnered considerable attention. This study aimed to evaluate the performance of these AI models in answering emergency medicine YDUS questions, marking the first assessment of the role of AI in this examination. The performance of 3 AI models (GPT-4, GPT-4o, and Gemini-Advanced) on questions from the emergency medicine YDUS examination was evaluated. The examination included 60 multiple-choice questions, of which 10% were publicly available. Questions were classified as clinical or factual. Responses of the AI models were analyzed using Cochran Q test as the omnibus test, with exact Bonferroni-adjusted McNemar tests for pairwise comparisons where applicable. No significant differences in the correct responses for both clinical and factual questions were observed between each AI model (P values: GPT-4o, 1.000; Gemini-Advanced, .554; and GPT-4, 1.000). GPT-4o significantly outperformed Gemini-Advanced in clinical (92.6% vs 70.4%) and factual questions (90.9% vs 78.8%) (P values: clinical, .021; factual, .039). A comparison of the overall performance showed an omnibus significant difference (P = .001); however, post hoc pairwise comparisons revealed that only GPT-4o (91.7%) significantly outperformed Gemini-Advanced (75%), whereas GPT-4 (88.3%) did not show a statistically significant difference from Gemini-Advanced after adjustment. This study found that both GPT-4o and GPT-4 significantly outperformed Gemini-Advanced in answering Turkish emergency medicine YDUS questions. While GPT-4o achieved the highest numerical accuracy, there was no statistically significant difference between GPT-4o and GPT-4. Both models demonstrated high accuracy in this examination dataset. Although these findings highlight their potential as supplementary learning tools, strong examination performance does not establish clinical readiness or definitive educational usefulness. Gemini-Advanced exhibited weaker performance but frequently advised expert consultation. However, this study did not formally evaluate ethical behavior, safety, or the appropriateness of these refusals.

نتیجه فارسی

این مطالعه عملکرد سه مدل هوش مصنوعی (GPT-4، GPT-4o و Gemini-Advanced) را در پاسخ به سوالات آزمون تخصصی پزشکی اورژانس ترکیه ارزیابی کرد. GPT-4o بالاترین دقت را داشت، در حالی که GPT-4 و GPT-4o اختلاف معناداری با یکدیگر نداشتند. Gemini-Advanced عملکرد ضعیف‌تری داشت.

  • ارزیابی عملکرد GPT-4، GPT-4o و Gemini-Advanced در آزمون YDUS پزشکی اورژانس ترکیه.
  • GPT-4o بالاترین دقت (۹۱.۷٪) را در پاسخ به سوالات داشت.
  • اختلاف معناداری آماری بین GPT-4o و GPT-4 وجود نداشت.
  • Gemini-Advanced عملکرد ضعیف‌تری داشت و به طور مکرر توصیه به مشاوره تخصصی می‌کرد.
  • عملکرد قوی در آزمون به تنهایی آمادگی بالینی یا کاربرد آموزشی قطعی را تأیید نمی‌کند.

ترجمه فارسی چکیده

این مطالعه به ارزیابی عملکرد سه مدل هوش مصنوعی (GPT-4، GPT-4o و Gemini-Advanced) در پاسخ به سوالات آزمون YDUS پزشکی اورژانس ترکیه پرداخت. آزمون شامل ۶۰ سوال چندگزینه‌ای بود که ۱۰٪ آن‌ها عمومی بودند. سوالات به دسته‌های بالینی و حقیقی طبقه‌بندی شدند. پاسخ‌ها با استفاده از آزمون Cochran Q و آزمون‌های دقیق Bonferroni-adjusted McNemar برای مقایسه‌های جفت‌به‌جفت تحلیل شدند. اختلاف معناداری در پاسخ‌های صحیح بین هر دو مدل هوش مصنوعی در سوالات بالینی و حقیقی مشاهده نشد (P-values: GPT-4o, 1.000; Gemini-Advanced, .554; و GPT-4, 1.000). GPT-4o به طور معناداری از Gemini-Advanced در سوالات بالینی (۹۲.۶٪ در مقابل ۷۰.۴٪) و سوالات حقیقی (۹۰.۹٪ در مقابل ۷۸.۸٪) عملکرد بهتری داشت (P-values: بالینی, .021; حقیقی, .039). مقایسه عملکرد کلی نشان‌دهنده اختلاف معنادار کلی بود (P = .001)، اما مقایسه‌های جفت‌به‌جفت نشان داد که تنها GPT-4o (۹۱.۷٪) به طور معناداری از Gemini-Advanced (۷۵٪) عملکرد بهتری داشت، در حالی که GPT-4 (۸۸.۳٪) پس از اصلاح، اختلاف معناداری با Gemini-Advanced نشان نداد. این مطالعه نشان داد که هر دو GPT-4o و GPT-4 به طور معناداری از Gemini-Advanced در پاسخ به سوالات پزشکی اورژانس ترکیه عملکرد بهتری داشتند. در حالی که GPT-4o بالاترین دقت عددی را کسب کرد، اختلاف معناداری آماری بین GPT-4o و GPT-4 وجود نداشت. هر دو مدل دقت بالایی در این مجموعه داده آزمون نشان دادند. اگرچه این یافته‌ها پتانسیل آن‌ها به عنوان ابزارهای یادگیری تکمیلی را برجسته می‌کند، عملکرد قوی در آزمون به تنهایی آمادگی بالینی یا کاربرد آموزشی قطعی را تأیید نمی‌کند. Gemini-Advanced عملکرد ضعیف‌تری داشت اما به طور مکرر مشاوره تخصصی توصیه کرد. با این حال، این مطالعه اخلاق‌شناسی، ایمنی یا مناسب بودن این امتناعات را به طور رسمی ارزیابی نکرد.

روش پژوهش

این مطالعه یک مطالعه ارزیابی‌کننده بود که عملکرد سه مدل هوش مصنوعی را بر اساس ۶۰ سوال چندگزینه‌ای از آزمون YDUS پزشکی اورژانس ترکیه بررسی کرد. تحلیل آماری با استفاده از آزمون Cochran Q و آزمون‌های دقیق Bonferroni-adjusted McNemar انجام شد.

محدودیت‌ها

این مطالعه اخلاق‌شناسی، ایمنی یا مناسبیت امتناعات هوش مصنوعی را به طور رسمی ارزیابی نکرد. همچنین عملکرد این مدل‌ها در سناریوهای بالینی واقعی یا در سایر آزمون‌ها بررسی نشد.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
سوالات آزمون YDUS پزشکی اورژانس ترکیه.
مداخله/مواجهه
سه مدل هوش مصنوعی: GPT-4، GPT-4o و Gemini-Advanced.
مقایسه
همه مدل‌ها با یکدیگر مقایسه شدند.
حجم نمونه
۶۰ سوال چندگزینه‌ای (۱۰٪ آن‌ها عمومی بود).

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

GPT-4oGemini-Advancedartificial intelligenceemergency medicinesubspecialty
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and enviro…

PubMed2026

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated respons…

PubMed2026

Liability and Standard of Care in AI-Driven Psychiatric Practice: European Viewpoint.

AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management…

PubMed2026

Routine laboratory panels classify internal medicine ICD-10 code groups: comparison with frontier large language models and laboratory-only specialist assessment.

INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from …