PubMed دسترسی آزاد

Can ChatGPT pass the polish national medical specialization examination in orthopedics and traumatology?

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

INTRODUCTION: Artificial intelligence (AI) has evolved rapidly in recent years and is becoming increasingly integrated into many areas of medicine. In the medical field, these systems have attracted considerable attention because of their potential applications in clinical decision support, medical education, scientific communication, and postgraduate training. The aim of the present study was to evaluate whether ChatGPT could achieve a passing score on the Polish National Medical Specialization Examination (Panstwowy Egzamin Specjalizacyjny, PES) in orthopedics and traumatology and to determine how different prompting strategies influenced its performance. MATERIALS AND METHODS: Authors systematically assessed the performance of ChatGPT-4 across five consecutive official PES exams (from Autumn 2023 to Spring 2025) in orthopedics and traumatology. Each exam was administered using three distinct prompting strategies: Professor prompt, Specialist prompt, Resident prompt. For each exam session (e.g., Spring 2024, Autumn 2023), the prompts were submitted to ChatGPT sequentially. The model's responses were evaluated against the official answer key published by the Polish Center of Medical Exams (Centrum Egzaminów Medycznych, CEM). RESULTS: The results were averaged across all prompts. The overall accuracy was 81% (range: 68.33%-85.83%). The highest score (85.83%) was recorded multiple times across different prompt types, indicating that the model could approach or exceed the minimum passing threshold depending on prompt structure. Agreement between prompting strategies was moderate to substantial (Cohen's κ 0.518-0.632; Fleiss' κ = 0.571). Radiology-related questions demonstrated the highest error rate among all question categories. CONCLUSIONS: Large language models such as ChatGPT can perform well on knowledge-based orthopedic examinations and may serve as useful tools for examination preparation and educational support. However, examination success should not be interpreted as evidence of clinical competence or readiness for independent surgical practice. Specialist certification and patient care remain dependent on human expertise, practical skills, and professional responsibility.

نتیجه فارسی

این مطالعه بررسی کرد که آیا مدل زبانی بزرگ چت‌جی‌پی‌تی-4 می‌تواند در آزمون تخصصی پزشکی لهستان در ارتوپدی و جراحی تراومات قبول شود. نتایج نشان داد که دقت کلی مدل ۸۱٪ بود و می‌توانست به نمره قبولی نزدیک شود، اما این موفقیت نباید به عنوان شواهدی از توانایی بالینی یا آمادگی برای عمل جراحی مستقل تفسیر شود.

  • چت‌جی‌پی‌تی-4 در آزمون‌های رسمی ارتوپدی و جراحی تراومات لهستان عملکردی معادل ۸۱٪ داشت.
  • استراتژی‌های مختلف درخواست (prompting) بر عملکرد مدل تأثیر داشتند.
  • سوالات مرتبط با رادیولوژی بالاترین نرخ خطا را داشتند.
  • این موفقیت به معنای توانایی بالینی یا آمادگی برای عمل جراحی مستقل نیست.
  • تخصص انسانی همچنان در گواهینامه تخصصی و مراقبت از بیمار ضروری است.

ترجمه فارسی چکیده

مقدمه: هوش مصنوعی (AI) در سال‌های اخیر به سرعت تکامل یافته و به تدریج در بسیاری از حوزه‌های پزشکی ادغام می‌شود. این سیستم‌ها در پزشکی به دلیل کاربردهای بالقوه در پشتیبانی از تصمیمات بالینی، آموزش پزشکی، ارتباط علمی و آموزش پس از فارغ‌التحصیلی توجه زیادی را به خود جلب کرده‌اند. هدف این مطالعه ارزیابی این بود که آیا چت‌جی‌پی‌تی می‌تواند نمره قبولی را در آزمون تخصصی پزشکی لهستان (PES) در ارتوپدی و جراحی تراومات کسب کند و اینکه چگونه استراتژی‌های مختلف درخواست (prompting) بر عملکرد آن تأثیر می‌گذارد. مواد و روش‌ها: نویسندگان عملکرد چت‌جی‌پی‌تی-4 را در پنج آزمون رسمی متوالی PES (از پاییز ۲۰۲۳ تا بهار ۲۰۲۵) در ارتوپدی و جراحی تراومات ارزیابی کردند. هر آزمون با سه استراتژی درخواست متفاوت اجرا شد: پروفسوری، تخصصی و کارآموزی. پاسخ‌های مدل با کلید پاسخ رسمی منتشر شده توسط مرکز آزمون‌های پزشکی لهستان (CEM) مقایسه شد. نتایج: میانگین نتایج در تمام درخواست‌ها محاسبه شد. دقت کلی ۸۱٪ (محدوده: ۶۸.۳۳٪ تا ۸۵.۸۳٪) بود. بالاترین نمره (۸۵.۸۳٪) چندین بار در انواع مختلف درخواست‌ها ثبت شد که نشان می‌دهد مدل می‌تواند به حداقل نمره قبولی نزدیک شود یا از آن فراتر رود بسته به ساختار درخواست. توافق بین استراتژی‌های درخواست متوسط تا قابل توجه بود (کوئین ۰.۵۱۸ تا ۰.۶۳۲؛ فلیس ۰.۵۷۱). سوالات مرتبط با رادیولوژی بالاترین نرخ خطا را در میان تمام دسته‌های سوال نشان داد. نتیجه‌گیری: مدل‌های زبانی بزرگ مانند چت‌جی‌پی‌تی می‌توانند در آزمون‌های دانش‌محور ارتوپدی عملکرد خوبی داشته باشند و می‌توانند به عنوان ابزارهای مفید برای آمادگی آزمون و پشتیبانی آموزشی عمل کنند. با این حال، موفقیت در آزمون نباید به عنوان شواهدی از توانایی بالینی یا آمادگی برای عمل جراحی مستقل تفسیر شود. گواهینامه تخصصی و مراقبت از بیمار همچنان به تخصص انسانی، مهارت‌های عملی و مسئولیت حرفه‌ای وابسته است.

روش پژوهش

نویسندگان عملکرد چت‌جی‌پی‌تی-4 را در پنج آزمون رسمی متوالی PES ارزیابی کردند. هر آزمون با سه استراتژی درخواست متفاوت اجرا شد و پاسخ‌ها با کلید پاسخ رسمی مقایسه شد.

محدودیت‌ها

محدودیت‌های گزارش نشده‌اند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
چت‌جی‌پی‌تی-4 (مدل زبانی بزرگ)
مداخله/مواجهه
استفاده از چت‌جی‌پی‌تی-4 در آزمون‌های تخصصی پزشکی لهستان در ارتوپدی و جراحی تراومات
مقایسه
کلید پاسخ رسمی آزمون‌های تخصصی پزشکی لهستان (CEM)
حجم نمونه
پنج آزمون رسمی متوالی (از پاییز ۲۰۲۳ تا بهار ۲۰۲۵)

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

AIChatGPTExamOrthopedicsTraumatology
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Near-Peer Anatomy-Anchored Teaching.

BACKGROUND: The transition from pre-clinical to clinical medicine is challenging, particularly in applying anatomical knowledge to patient care. Reductions in dedicated anatomy teaching time have compounded this. Near-peer teaching may help address this gap by reducing hierarchy and enhancing psychological safety, though few programmes have explicitly targeted the pre-clinical to clinical transition through the integration of anatomy w…

PubMed2026

Additively manufactured PLA-bioceramic porous orthopedic implants for biomedical applications: from material-structure co-design to clinical translation.

Porous orthopedic implants are widely studied because they promote osseointegration and reduce stress shielding associated with dense, stiff implants. This review critically examines recent progress in the design and fabrication of additively manufactured porous orthopedic implants, with a focused emphasis on polylactic acid (PLA)-bioceramic systems, particularly hydroxyapatite-, tricalcium phosphate-, and bioactive glass-containing co…

PubMed2026

Implementation of evidence-based practice guidelines to prevent urinary retention in hospitals: A process evaluation in orthopaedic care.

BACKGROUND AND OBJECTIVES: National clinical practice guidelines promote evidence-based practice, though are hardly implemented without local facilitation. While further research is needed as for what facilitation strategies work, in what context, and with what outcomes, the Onset PrevenTIon of urinary retention in Orthopaedic Nursing and rehabilitation (OPTION) trialled a tailored implementation intervention for evidence-based clinica…

PubMed2026

Advance Care Planning in Orthopedics.

Given that about two-thirds of adults in the United States have not completed any advance directive, and there are over five million inpatient and outpatient orthopedic surgeries each year, there are unparalleled opportunities to engage with patients in advance care planning (ACP) discussions. ACP is defined as an ongoing strategy to help adults across the lifespan identify and communicate their preferences for medical care, based on t…