PubMed چکیده/رکورد

Performance of ChatGPT, Claude, and AMBOSS on the European Board of Urology In-Service Assessment and Alignment With the European Association of Urology 2025 Guidelines: Comparative Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Recent advances in AI, particularly large language models, have generated growing interest in their application to medical education and examination preparation. However, the accuracy, reasoning quality, and adherence to clinical guidelines of these tools in postgraduate urology assessments remain unclear. OBJECTIVE: This study aimed to evaluate the performance of 3 AI tools, ChatGPT (GPT-4.0), Claude (version 4.5), and AMBOSS, on European Board of Urology (EBU)-style multiple-choice questions, with a particular focus on accuracy, insight, concordance, and adherence to European Association of Urology (EAU) guidelines. METHODS: A total of 200 single-best-answer questions from the EBU In-Service Assessment workbook (2021-2022) were input into each AI model. Models were prompted to select an answer and provide an explanation. Two urologists with post-Fellowship of the Royal College of Surgeons (FRCS) training independently assessed the outputs. Accuracy was defined as correct answer selection. Concordance was defined as the logical alignment between the answer and its explanation. Insight was evaluated across 3 domains-nonobvious deduction, discriminative reasoning, and clinical validity-and was graded as low, moderate, or high. RESULTS: ChatGPT demonstrated the highest accuracy (171/200, 85.5%), compared to Claude and AMBOSS (both 159/200, 79.5%; P=.14). Concordance was also significantly higher for ChatGPT (190/200, 95%) than for Claude (176/200, 88%) and AMBOSS (152/200, 76%; P<.001). Nonobvious deduction was predominantly low to moderate across all models, reflecting the recall-based nature of many questions. ChatGPT and Claude showed stronger discriminative reasoning, while AMBOSS demonstrated limited exclusion of alternative options. Clinical validity was high overall, with ChatGPT showing the greatest consistency with EAU guidelines. There was substantial agreement between the 2 reviewers (weighted κ coefficient >0.61). CONCLUSIONS: AI tools can achieve high accuracy on EBU-style assessments; however, differences in reasoning quality and guideline adherence are evident. ChatGPT demonstrated superior performance across all evaluated domains, supporting its role as a potential adjunct in postgraduate urology education.

نتیجه فارسی

این مطالعه مقایسه‌ای عملکرد سه مدل هوش مصنوعی (ChatGPT، Claude و AMBOSS) را در پاسخ به سوالات امتحانی اورولوژی اروپا بررسی کرد. نتایج نشان داد که ChatGPT بالاترین دقت و همسویی را با دستورالعمل‌های EAU داشت. تفاوت‌هایی در کیفیت استدلال و رعایت دستورالعمل‌ها بین ابزارها وجود داشت.

  • ChatGPT بالاترین دقت (۸۵.۵٪) و همسویی (۹۵٪) را نسبت به Claude و AMBOSS نشان داد.
  • استدلال غیرمستقیم در تمام مدل‌ها عمدتاً کم تا متوسط بود.
  • اعتبار بالینی به طور کلی بالا بود و ChatGPT بیشترین ثبات را با دستورالعمل‌های EAU نشان داد.
  • دو متخصص اورولوژی توافق قابل توجهی در ارزیابی خروجی‌ها داشتند.

ترجمه فارسی چکیده

پسرفت‌های اخیر هوش مصنوعی، به‌ویژه مدل‌های زبانی بزرگ، توجه زیادی را به کاربرد آن‌ها در آموزش پزشکی و آمادگی برای امتحانات جلب کرده است. با این حال، دقت، کیفیت استدلال و رعایت دستورالعمل‌های بالینی این ابزارها در ارزیابی‌های پس‌دکتری اورولوژی همچنان نامشخص است. این مطالعه به ارزیابی عملکرد سه ابزار هوش مصنوعی، ChatGPT (نسخه 4.0)، Claude (نسخه 4.5) و AMBOSS بر روی سوالات چندگزینه‌ای سبک کانون اورولوژی اروپا (EBU) با تمرکز بر دقت، بینش، همسویی و رعایت دستورالعمل‌های انجمن اورولوژی اروپا (EAU) پرداخت. مجموعه‌ای از ۲۰۰ سوال بهترین پاسخ تک‌گزینه‌ای از کتابچه‌کار ارزیابی داخلی کانون اورولوژی اروپا (۲۰۲۱-۲۰۲۲) به هر یک از مدل‌های هوش مصنوعی وارد شد. مدل‌ها دستورالعمل دریافت کردند که یک پاسخ انتخاب کرده و توضیحی ارائه دهند. دو متخصص اورولوژی با آموزش پس از دریافت مدرک فوق‌تخصص جراحی‌های سلطنتی (FRCS) خروجی‌ها را به طور مستقل ارزیابی کردند. دقت به عنوان انتخاب پاسخ صحیح تعریف شد. همسویی به عنوان همسویی منطقی بین پاسخ و توضیح آن تعریف شد. بینش در سه حوزه - استدلال غیرمستقیم، استدلال تمایز‌بخش و اعتبار بالینی - ارزیابی شد و به صورت کم، متوسط یا بالا رتبه‌بندی شد. ChatGPT بالاترین دقت (۱۷۱/۲۰۰، ۸۵.۵٪) را نشان داد، در حالی که Claude و AMBOSS هر دو ۱۵۹/۲۰۰ (۷۹.۵٪) داشتند (P=.14). همسویی برای ChatGPT نیز به طور معناداری بالاتر (۱۹۰/۲۰۰، ۹۵٪) نسبت به Claude (۱۷۶/۲۰۰، ۸۸٪) و AMBOSS (۱۵۲/۲۰۰، ۷۶٪) بود (P<.001). استدلال غیرمستقیم در تمام مدل‌ها عمدتاً کم تا متوسط بود که بازتاب‌دهنده ماهیت مبتنی بر حافظه بسیاری از سوالات است. ChatGPT و Claude نشان‌دهنده استدلال تمایزبخش قوی‌تری بودند، در حالی که AMBOSS نشان داد محدودیت در حذف گزینه‌های جایگزین. اعتبار بالینی به طور کلی بالا بود، با اینکه ChatGPT بیشترین ثبات را با دستورالعمل‌های EAU نشان داد. بین دو بازبینی‌کننده توافق قابل توجهی وجود داشت (ضریب κ وزنی >0.61). ابزارهای هوش مصنوعی می‌توانند دقت بالایی در ارزیابی‌های سبک EBU کسب کنند؛ با این حال، تفاوت‌ها در کیفیت استدلال و رعایت دستورالعمل‌ها آشکار است. ChatGPT عملکرد برتر را در تمام حوزه‌های ارزیابی شده نشان داد که نقش آن به عنوان یک مکمل بالقوه در آموزش اورولوژی پس‌دکتری را پشتیبانی می‌کند.

روش پژوهش

۲۰۰ سوال از کتابچه‌کار ارزیابی داخلی کانون اورولوژی اروپا (۲۰۲۱-۲۰۲۲) به هر سه مدل هوش مصنوعی وارد شد. دو متخصص اورولوژی با آموزش FRCS خروجی‌ها را به طور مستقل ارزیابی کردند.

محدودیت‌ها

محدودیت‌های گزارش نشده‌اند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
سوالات امتحانی اورولوژی سبک کانون اورولوژی اروپا (EBU).
مداخله/مواجهه
مدل‌های هوش مصنوعی ChatGPT (نسخه 4.0)، Claude (نسخه 4.5) و AMBOSS.
مقایسه
همه یکسان (بررسی مقایسه‌ای).
حجم نمونه
۲۰۰ سوال.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

AIAMBOSSChatGPTClaudeEAU guidelinesEBUEuropean Association of Urology guidelinesEuropean Board of Urologyartificial intelligence
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Institutional availability, training resources, and supervised participation in robot-assisted surgery among Polish urology residents: a nationwide survey.

Robot-assisted surgery (RAS) has expanded rapidly in Poland, but its integration into resident education remains uncertain. In a nationwide anonymous 35-question survey sent to 390 Polish urology residents, 156 completed it (40.0%). Among respondents, 110 (70.5%) reported RAS availability in their training centre; 56 (35.9%) reported robotic training access, 43 (27.6%) a robotic course, 46 (29.5%) simulator availability, 76 (48.7%) ass…

PubMed2026

[The wife of the Nobel Laurate: Elsbet Forßmann née Engel (1906-1993)].

This article highlights the life and work of the urologist Elsbet Forßmann (1906-1993), who was remembered for a very long time as the "brave and understanding wife" of Nobel Prize laureate Werner Forßmann (1904-1979) and draws attention to the widespread invisibility of women in the culture of memory of the field. But she was far more than that. After completing her training under Karl Heusch (1894-1986) at the Rudolf Virchow Hospital…

PubMed2026

Exploring the impact of artificial intelligence on radiation dose reduction in urological imaging: a systematic review from EAU endourology.

BACKGROUND: Patients with urological conditions often undergo recurrent computed tomography (CT) imaging, which results in cumulative radiation exposure which can be deleterious. Artificial intelligence (AI)- based technologies, including Deep Learning Image Reconstruction (DLIR), have emerged as a potential strategy to reduce radiation dose in CT imaging while maintaining image quality. This systematic review aimed to quantify the ben…

PubMed2026

Worldwide variations in purchase and per-procedure costs of single-use flexible ureteroscopes and flexible navigable access sheaths: the economic impact of reuse strategies on procedural outcomes from a 40-country YAU-EAU endourology international collaboration.

Endourological practice increasingly relies on single-use devices, the costs of which vary considerably across regions and healthcare systems. Flexible single-use ureteroscopes (FURS) and flexible and navigable suction ureteral access sheaths (FANS) represent a substantial proportion of ureteroscopy-related expenditure. To mitigate these costs, some centres have adopted reuse strategies aimed at reducing per-procedure expenses. The aim…